Pith. sign in

REVIEW 4 major objections 6 minor 264 references

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

T0 review · 4 major / 6 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Hallo4D claims that spatiotemporal hallucinations in 3D and 4D generation can be detected by large multimodal language models and corrected through consensus-driven, image-space editing, without retraining or architectural changes.

desk verdict A useful extension of Hallo3D to 4D, but the core correction loss depends on an untested assumption and the 'no architectural changes' claim is contradicted by the paper's own attention replacement. read the letter →

arxiv 2607.12752 v2 pith:VYAQSVCI submitted 2026-07-14 cs.CV cs.AI

classification cs.CVcs.AI
keywords 3Dgeneration4Dhallucinationmitigationlargemultimodallanguagemodelsscoredistillationsamplingspatiotemporalconsistencyimage-spaceeditingconsensusvoting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Hallo4D argues that the hallucinations plaguing 3D and 4D generation—duplicated heads, missing geometry, identity flicker, jitter—can be caught after the fact and corrected without retraining or changing the underlying model. Its central claim is that large multimodal language models can serve as both detectors, turning rendered views and frames into targeted negative prompts, and as judges, picking the most geometrically faithful candidate edit through multi-model voting. The corrected image then enters the optimization loop as an image-space loss derived from Score Distillation Sampling. A sympathetic reader would care because the method is plug-and-play across many existing generation pipelines, treating inconsistency as a detectable, correctable signal rather than something baked into training. The paper reports consistent gains over its baselines across text-to-3D, image-to-3D, and 4D generation.

What carries the argument

The load-bearing mechanism is the generation-detection-correction loop. A large multimodal language model inspects multi-view or multi-frame renderings and returns standardized inconsistency reports plus enhanced negative prompts; a diffusion-based image editor (DDIM inversion followed by sampling under those negative prompts) produces several candidate corrected images; and a second language-model ensemble scores those candidates and selects one by majority vote. That selected image feeds an image-space consistency loss whose legitimacy is argued by rewriting SDS in image-prediction form. Supporting machinery includes appearance attention that borrows key and value features from a focal vie

What would settle it

Take a 3D model with a known duplicated feature, run Hallo4D's detector and edit, and check in novel views whether the chosen candidate actually removes that feature; then intentionally feed a wrong negative prompt and measure whether the consistency loss degrades geometry more than plain SDS, which would show the loop trusts the edit too much. A more direct test: on a held-out set of multi-view renderings, count how often the edited image is judged more view-consistent than the diffusion-clean prediction by independent human raters or a separate geometric check, not the same language-model en

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that spatiotemporal inconsistency in 3D and 4D generation has a tractable image-space correction loop. Hallo4D renders a view or frame, asks a large multimodal language model to identify what is wrong across views or timesteps and to phrase the failure as a negative prompt, regenerates the image with DDIM inversion guided by that prompt, and then lets an ensemble of language models vote among several candidate regenerations to pick the one that best preserves structure while removing the artifact. The selected image becomes the target of a consistency loss applied alongside SDS. The authors further derive an image-space reformulation of SDS (Eq. 9–1

Load-bearing premise

The load-bearing premise is that the image regenerated by editing the rendered view under the language model's negative prompt is a more faithful geometric target than the diffusion model's own clean prediction; if that edit is sometimes less faithful, the correction loss could reinforce new hallucinations instead of removing old ones.

Editorial extensions

If this is right

  • If a generator's rendered views are semantically complete, Hallo4D can be overlaid on it to reduce multi-view artifacts such as duplicated or missing structures without touching its weights.
  • For 4D pipelines, early language-model-guided initialization and flow-based keyframe sampling should reduce temporal drift, identity flicker, and jitter relative to uniform frame sampling.
  • The exposure-aware losses plus union-of-frusta pruning should prevent white-out and black-out collapse under non-frontal cameras, a failure mode standard SDS leaves largely unaddressed.
  • Because corrections happen in image space and are consensus-selected, the framework should avoid compounding single-edit errors across optimization iterations.
  • The image-space reformulation of SDS implies that pixel-space consistency losses are compatible with SDS training dynamics, so the wrapper should extend to other SDS-based generators.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the detector-as-judge split is an implicit scaling argument—as language models improve at spatial and temporal reasoning, this wrapper should inherit those gains with no change to the 3D or 4D model, a claim testable by swapping the language-model backbone.
  • Beyond the paper: the constrained evaluation (single foreground subject, fixed-orbit cameras) leaves multi-object scenes and free camera trajectories untested; extending the framework would require occlusion-aware detection and voting.
  • Beyond the paper: the framework is only as honest as its judges; replacing or augmenting language-model voting with an automatic geometric check, such as silhouette or multi-view feature consistency, would decouple correction quality from the judge's bias.
  • Beyond the paper: because the correction target is produced by a 2D diffusion edit, the same 2D prior that caused the hallucination may bias the target; blending the edited image with the original or limiting the edit to early denoising steps is a tunable lever the paper does not calibrate.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. Hallo4D proposes a model-agnostic, plug-in framework for reducing spatial and temporal hallucinations in 3D and 4D generation. It introduces a generation-detection-correction loop in which LMMs analyze multi-view/multi-frame renderings, produce negative prompts, and guide a consensus-driven image-space re-consistency loss (LCG) based on DDIM inversion/regeneration with LMM majority voting. Additional components include a multi-view appearance attention module (AAttn), LMM-guided initialization for 4D, optical-flow-based keyframe sampling (OF-Range), exposure-aware losses (CSEA, LDR), and union-of-frusta visibility pruning. The paper claims consistent improvements over several text-to-3D, image-to-3D, and 4D baselines, supported by qualitative comparisons, CLIP-based metrics, and user studies.

Significance. If the central premise is validated, Hallo4D is a potentially significant contribution: it addresses a real failure mode (spatiotemporal hallucination) in a largely model-agnostic way and is evaluated across representative 3D and 4D pipelines. The image-space reformulation of SDS in Eqs. (9)–(10) is a standard and clearly presented algebraic identity, and the paper is honest about the conditional nature of the correction loss (“provided that \x e4̂h_0 exhibits stronger view consistency…”). The breadth of baselines, the user study, and the modular ablations are strengths. However, the load-bearing premise that the consensus-edited image is a more faithful target than the diffusion model’s own prediction is not directly established, and the quantitative claims are reported without variance or significance information. The contradiction between the “no architectural modifications” claim and the AAttn module also needs resolution. These issues are fixable, but the current evidence is conditional.

major comments (4)
  1. [Sec. 4.3, Eq. (19)] The LCG loss is theoretically justified only if the consensus-edited image \x e4̂h*_0 is a more faithful target than the SDS clean prediction x^φ_0. The paper states this premise explicitly but supports it with a single attention visualization (Fig. 5), which is not a fidelity measurement. It does not show that the edited images are geometrically more correct, nor that the LMM selector’s votes track ground-truth fidelity. If the edit over-smooths or introduces new artifacts, LCG will pull the 3D/4D representation toward an unreliable target and can amplify hallucinations. Please add a quantitative validation: e.g., on renderings with known ground-truth geometry, compare errors (LPIPS, geometric IoU, view-consistency metrics) of \x e4̂h*_0 versus x^φ_0, and report the selector’s agreement with ground-truth fidelity. This is the load-bearing point for the proposed loss.
  2. [Abstract, Sec. 1, Sec. 4.1, Eq. (5)] The claim that Hallo4D works “without requiring retraining or architectural modifications” is contradicted by Sec. 4.1, where AAttn replaces the original self-attention in the U-Net denoiser with a cross-view attention mechanism. This is an architectural modification of the 2D diffusion backbone, even if the 3D/4D representation itself is unchanged. Please either remove or carefully qualify the “no architectural modifications” claim throughout the paper, or explicitly distinguish “no modification to the 3D/4D representation” from “no modification to the diffusion U-Net.”
  3. [Tables 1–3, Sec. 5.3] All quantitative results are reported as single point estimates without error bars, seeds, or significance tests. Several differences are very small (e.g., Table 2, DreamGaussian: CD 0.0185 vs 0.0185, PSNR 16.502 vs 16.530, SSIM +0.026), while user-study differences are much larger. Without variance information or paired statistical tests, the reader cannot judge whether the improvements are reliable. Please report mean±std over seeds/objects and appropriate significance tests, and consider releasing code to enable independent verification. The user study should also report inter-subject variability or confidence intervals.
  4. [Sec. 5.4, Table 4] The ablation study does not isolate the specific premise of the consensus selector. The “w/o Re-Cons.” ablation removes the entire correction mechanism, and the “w/o Consensus” row in Table 4 is incremental for 4D only; the text notes that the baseline excluding CESA & LDR “inherently omits the Consensus module as well.” Thus there is no ablation that compares the consensus-edited target to a single random edit or to x^φ_0 while keeping the same detection pipeline. Please add such an ablation; this directly tests the paper’s core claim that multi-model voting prevents compounding errors.
minor comments (6)
  1. [Sec. 4, first paragraph] The sentence “Across both 3D and 4D stages, identified inconsistencies are corrected through a Prompt-Enhanced Re-consistency module…” is repeated nearly verbatim in consecutive paragraphs. Please remove the duplicate.
  2. [Eq. (19)] The expectation in LCG is not fully specified. Please define the distribution over rendered views, cameras, and timesteps over which the expectation is taken.
  3. [Sec. 5.1, Eq. (23)] The text says “we apply a decaying weight to Rτ in Eq. (23)” but Eq. (23) contains no Rτ; the decaying weight presumably applies to γ. Please correct the variable name.
  4. [Table 4, Sec. 5.4] The module name is written inconsistently as both “CESA” and “CSEA.” Please unify the abbreviation.
  5. [References] References Xu et al. 2024b and 2024c refer to the same paper (MagicAnimate) and should be merged or renumbered. Also check duplicate entries in the reference list.
  6. [Sec. 5.1, focal view selection] The criterion “first view where Fovy exceeds 120% of the baseline default” is vague. Please specify the actual camera configurations and threshold used.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: Eq. 9 is an algebraic identity, and the LCG target premise is an empirical assumption rather than a fitted input.

full rationale

The derivation chain is self-contained. Eq. 9 follows by substituting Eq. 1 into Eq. 4 and is a standard algebraic identity, not a fitted relation. Eq. 19 defines a new image-space MSE target, and its usefulness is explicitly conditional on an empirical premise: 'provided that xhat_0 exhibits stronger view consistency and better alignment with realistic semantics than x_phi_0.' The paper supports that premise with a qualitative attention visualization (Fig. 5); this is evidence weakness, not a construction-level circularity, because xhat_0 is produced by DDIM inversion plus LMM consensus voting rather than fitted to the evaluation metrics or to x_0 in a way that would force the reported gains. The only self-citation (Hallo3D, Wang et al. 2024a) appears in the 'Differences from Our Prior Work' passage as a scope statement and is not used to justify the SDS reformulation or the consensus loss; theoretical support instead cites external works (Kingma et al. 2021; Deng et al. 2024; Song et al. 2022). No uniqueness theorem or ansatz is imported from the authors' prior work. External comparisons against unmodified baselines, objective metrics, and user studies provide independent content. The w/o Re-Cons. ablation removes the whole correction and does not isolate the comparative premise, but that is a limitation in evidence, not a circular reduction. Therefore no circular step is exhibited.

Assumptions & free parameters 8 free parameters · 6 assumptions · 0 invented entities

The central claim rests on LMM capability assumptions, a target-improvement assumption, and several empirically set hyperparameters. No new physical entities are introduced.

free parameters (8)
  • w1 (LCG weight in Eq 20) = 0.1
    Balances image-space correction loss against SDS; set in Sec 5.1.
  • w2 (L_CG-init weight in Eq 21) = decays 0.1 -> 0.01
    Weights initialization correction during the first 10% of training; empirical decay schedule.
  • gamma (regularization coefficient in Eq 23) = decays 1 -> 0.0001
    Controls penalty on sampling near existing keyframes; tuned to balance diversity.
  • mu (LSE temperature in Eq 26) = 0.06
    Controls smooth maximum of exposure penalties in CSEA; chosen empirically.
  • sigma_0 (target log-luminance std in Eq 30) = 0.9
    Desired contrast bandwidth for the LDR loss; chosen empirically.
  • vartheta (log stabilizer in Eq 30) = 0.01
    Prevents singular gradients in dark regions.
  • focal view selection threshold = 120% of baseline Fovy
    Selects the focal view for appearance attention; heuristic.
  • candidate count n and evaluator count m = n=4, m=4
    Number of candidate edits and LMM voters used in the consensus selector.
assumptions (6)
  • standard math Image-space SDS reformulation (Eq 9) is valid and can be combined with the LCG loss.
    The algebra is shown in Sec 4.3, but its use as a training loss depends on the next assumption.
  • domain assumption The LMM detector reliably localizes spatial and temporal inconsistencies in constrained renderings.
    Sec 4.2 assumes LLaVA-OneVision-72B's cross-view reasoning transfers to unseen renderings; only qualitative evidence is given in Fig 4.
  • domain assumption The edited image x-hat_0 is a better optimization target than the SDS clean prediction x_phi_0.
    Sec 4.3 states effectiveness 'provided that x-hat_0 exhibits stronger view consistency'; this is not proven and is supported only by attention visualization.
  • domain assumption Replacing U-Net self-attention with focal-view cross-attention does not destabilize denoising.
    Sec 4.1 modifies the diffusion U-Net architecture without retraining; no theoretical guarantee is provided.
  • domain assumption CLIP-based exposure prompts and the LDR loss drive correct exposure semantics.
    Sec 4.6 relies on CLIP text-image alignment for exposure judgments and on the chosen log-luminance target.
  • domain assumption Optical-flow saliency selects frames whose refinement improves temporal consistency.
    Sec 4.5 assumes motion saliency equals optimization importance for 4D training.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation." pith.science (2026). https://pith.science/paper/VYAQSVCI

@misc{pith2026260712752,
  author       = {Pith},
  title        = {Pith review of: Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VYAQSVCI}},
  note         = {Machine review of arXiv:2607.12752}
}
read the original abstract

While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry. These issues become more severe in 4D generation, where maintaining consistency across viewpoints and temporal evolution introduces additional challenges, including jitter, identity flicker, and structural drift. We present \textbf{Hallo4D}, a unified and model-agnostic framework for mitigating spatiotemporal hallucinations in 3D and 4D content generation. Hallo4D introduces a generation-detection-correction paradigm that leverages large multimodal language models (LMMs) to identify and summarize spatial and temporal inconsistencies from multi-view and multi-frame renderings. These insights guide a consensus-driven image-space consistency optimization, where an LMM-based selector evaluates candidate corrections through multi-model voting, without requiring retraining or architectural modifications. To further improve temporal consistency and optimization efficiency, Hallo4D incorporates motion-aware keyframe sampling, LMM-guided initialization, and appearance alignment. We additionally introduce exposure-aware optimization and visibility pruning to enhance robustness under challenging viewpoints. Extensive experiments demonstrate that Hallo4D consistently outperforms strong baselines across diverse 3D and 4D generation settings, providing a scalable and generalizable solution for consistency-aware content generation.

Figures

Figures reproduced from arXiv: 2607.12752 by the authors.

Figure 1
Figure 1. Comparison of 3D and 4D generation results between Hallo4D and the baseline. Our method achieves significant improvements in spatiotemporal consistency. leads to hallucinated or duplicated structures in unobserved views—a failure mode commonly referred to as the Janus problem(Armandpour et al., 2023). This issue reflects a core weakness of current pipelines: the absence of mech￾anisms to explicitly model and enforce… view at source ↗
Figure 2
Figure 2. Pipeline overview of the 3D consistency optimization stage. We jointly optimize our model using LSDS and LCG. For LSDS, a focal view is selected based on camera pose and used as keys and values to align all views via attention. For LCG, hallucinations are identified using LMMs and translated into enhanced negative prompts. These prompts guide the generation of multiple candidate corrections, from which an MLLM conse… view at source ↗
Figure 3
Figure 3. Pipeline overview of the 4D spatiotemporal consistency stage. We first identify the views and frames with the most severe inconsistencies through initialization to guide early-stage training. During training, we employ OF-Range, an optical flow-based inter-frame sampling strategy that improves both the efficiency and effectiveness of the 4D generation process. Multi-view Appearance Alignment strategy by introducing … view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: A multi-modal reasoning case study designed to evaluate the capabilities of LMMs in spatial structure reasoning, multi-view inconsistency diagnosis, and temporal inconsistency localization. The first round of dialogue demonstrates that LMMs possess the necessary reason…
Figure 5
Figure 5. Figure 5: Attention Visualization in Prompt-Enhanced Re￾consistency. Left: image shows an inconsistent rendered view. Middle and Right: images visualize U-Net attention from the DDIM-Inversion-based image editing model, with the middle and rightmost images corresponding to early…
Figure 6
Figure 6. Figure 6: Illustration of over-exposure and under-exposure situation. Baselines exhibit view-dependent collapse in non-frontal views, producing white-out/black-out frames, whereas Hallo4D, with losses of CSEA and LDR, maintains stable exposure and preserves structure across view…
Figure 7
Figure 7. Figure 7: Qualitative comparison in text-driven 3D generation of Hallo4D and baseline models. For intuitive comparison, we use the same and complementary viewpoints to present the results, where our method shows significant improvements. combine signals from multiple exposure-re…
Figure 8
Figure 8. Figure 8: Qualitative comparison in image-driven 3D generation of Hallo4D and baseline models. For intuitive compar￾ison, we enlarge the regions with noticeable differences, and the results demonstrate that our method achieves significant improvements. Moreover, we regularize th…
Figure 9
Figure 9. Figure 9: Qualitative comparison in 4D generation of Hallo4D and baseline models. Among them, Consistent4D and DreamGaussian4D are image-to-4D methods, while 4D-FY is a text-to-4D generation approach. due to the presence of LCG-init, we apply a decaying weight to Rτ in Eq. (23),…
Figure 10
Figure 10. Figure 10: Ablation study of Hallo4D. We conduct a quantitative ablation study by successively removing each module from the full model, confirming the necessity and effectiveness of the modular design in achieving high-quality spatiotemporal generation [PITH_FULL_IMAGE:figures…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

264 extracted references · 113 linked inside Pith

  1. [1]

    Azuma, Daichi and Miyanishi, Taiki and Kurita, Shuhei and Kawanabe, Motoaki , year =

  2. [2]

    IJCV , author =

    A. IJCV , author =. 2026 , eprint =

  3. [3]

    2503.20785 , archiveprefix =

    Liu, Tianqi and Huang, Zihao and Chen, Zhaoxi and Wang, Guangcong and Hu, Shoukang and Shen, Liao and Sun, Huiqiang and Cao, Zhiguo and Li, Wei and Liu, Ziwei , year =. 2503.20785 , archiveprefix =

  4. [4]

    2025 , eprint =

    arXiv:2508.13154 , author =. 2025 , eprint =

  5. [5]

    arXiv:2503.14501 , author =

    Advances in. arXiv:2503.14501 , author =. 2025 , eprint =

  6. [6]

    and Vora, Sourabh and Liong, Venice Erin and Xu, Qiang and Krishnan, Anush and Pan, Yu and Baldan, Giancarlo and Beijbom, Oscar , year =

    Caesar, Holger and Bankiti, Varun and Lang, Alex H. and Vora, Sourabh and Liong, Venice Erin and Xu, Qiang and Krishnan, Anush and Pan, Yu and Baldan, Giancarlo and Beijbom, Oscar , year =

  7. [7]

    and Savva, Manolis and Halber, Maciej and Funkhouser, Thomas and Nie

    Dai, Angela and Chang, Angel X. and Savva, Manolis and Halber, Maciej and Funkhouser, Thomas and Nie

  8. [8]

    Hosseinzadeh, Mehrdad and Wang, Yang , year =. Image

Show all 264 references
  1. [9]

    Learning to

    Jhamtani, Harsh and. Learning to

  2. [10]

    Shridhar, Mohit and Thomason, Jesse and Gordon, Daniel and Bisk, Yonatan and Han, Winson and Mottaghi, Roozbeh and Zettlemoyer, Luke and Fox, Dieter , year =

  3. [11]

    Suhr, Alane and Zhou, Stephanie and Zhang, Ally and Zhang, Iris and Bai, Huajun and Artzi, Yoav , year =. A

  4. [12]

    Brown, Tom B. and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and Agarwal, Sandhini and. Language

  5. [13]

    AlBahar, Badour and Saito, Shunsuke and Tseng, Hung-Yu and Kim, Changil and Kopf, Johannes and Huang, Jia-Bin , year =. Single-

  6. [14]

    Re-Imagine the

    Armandpour, Mohammadreza and Sadeghian, Ali and Zheng, Huangjie and Sadeghian, Amir and Zhou, Mingyuan , year =. Re-Imagine the. arXiv:2304.04968 , eprint =

  7. [15]

    and Ovsjanikov, Maks , year =

    Attaiki, Souhaib and Guerrero, Paul and Ceylan, Duygu and Mitra, Niloy J. and Ovsjanikov, Maks , year =. 2412.16717 , archiveprefix =

  8. [16]

    , year =

    Bahmani, Sherwin and Skorokhodov, Ivan and Rong, Victor and Wetzstein, Gordon and Guibas, Leonidas and Wonka, Peter and Tulyakov, Sergey and Park, Jeong Joon and Tagliasacchi, Andrea and Lindell, David B. , year =

  9. [17]

    , year =

    Bahmani, Sherwin and Liu, Xian and Yifan, Wang and Skorokhodov, Ivan and Rong, Victor and Liu, Ziwei and Liu, Xihui and Park, Jeong Joon and Tulyakov, Sergey and Wetzstein, Gordon and Tagliasacchi, Andrea and Lindell, David B. , year =. 2403.17920 , archiveprefix =

  10. [18]

    arXiv:2211.01324 , eprint =

    Balaji, Yogesh and Nah, Seungjun and Huang, Xun and Vahdat, Arash and Song, Jiaming and Kreis, Karsten and Aittala, Miika and Aila, Timo and Laine, Samuli and Catanzaro, Bryan and Karras, Tero and Liu, Ming-Yu , year =. arXiv:2211.01324 , eprint =

  11. [19]

    and Mildenhall, Ben and Verbin, Dor and Srinivasan, Pratul P

    Barron, Jonathan T. and Mildenhall, Ben and Verbin, Dor and Srinivasan, Pratul P. and Hedman, Peter , year =. Mip-

  12. [20]

    Gaussian Process-Based Representation Learning via Timeseries Symmetries , booktitle =

    Bevanda, Petar and Beier, Max and Lederer, Armin and Capone, Alexandre and Sosnowski, Stefan and Hirche, Sandra , year =. Gaussian Process-Based Representation Learning via Timeseries Symmetries , booktitle =

  13. [21]

    Training

    Black, Kevin and Janner, Michael and Du, Yilun and Kostrikov, Ilya and Levine, Sergey , year =. Training

  14. [22]

    Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets , shorttitle =

    Blattmann, Andreas and Dockhorn, Tim and Kulal, Sumith and Mendelevitch, Daniel and Kilian, Maciej and Lorenz, Dominik and Levi, Yam and English, Zion and Voleti, Vikram and Letts, Adam and Jampani, Varun and Rombach, Robin , year =. Stable Video Diffusion: Scaling Latent Vide...

  15. [23]

    IEEE Trans

    Cao, Ziang and Hong, Fangzhou and Wu, Tong and Pan, Liang and Liu, Ziwei , year =. IEEE Trans. Pattern Anal. Mach. Intell. , volume =

  16. [24]

    Cao, Ang and Johnson, Justin , year =

  17. [25]

    2304.08465 , archiveprefix =

    Cao, Mingdeng and Wang, Xintao and Qi, Zhongang and Shan, Ying and Qie, Xiaohu and Zheng, Yinqiang , year =. 2304.08465 , archiveprefix =

  18. [26]

    2404.18929 , archiveprefix =

    Chen, Minghao and Laina, Iro and Vedaldi, Andrea , year =. 2404.18929 , archiveprefix =

  19. [27]

    Chen, Rui and Chen, Yongwei and Jiao, Ningxin and Jia, Kui , year =

  20. [28]

    2404.06425 , archiveprefix =

    Cheng, Ta-Ying and Sharma, Prafull and Markham, Andrew and Trigoni, Niki and Jampani, Varun , year =. 2404.06425 , archiveprefix =

  21. [29]

    2403.15951 , archiveprefix =

    Chen, Jiacheng and Wu, Yuefan and Tan, Jiaqi and Ma, Hang and Furukawa, Yasutaka , year =. 2403.15951 , archiveprefix =

  22. [30]

    IEEE Trans

    Chen, Zhaoxi and Wang, Guangcong and Liu, Ziwei , year =. IEEE Trans. Pattern Anal. Mach. Intell. , volume =

  23. [31]

    Chen, Cheng and Yang, Xiaofeng and Yang, Fan and Feng, Chengzeng and Fu, Zhoujie and Foo, Chuan-Sheng and Lin, Guosheng and Liu, Fayao , year =

  24. [32]

    A Survey on

    Chen, Guikun and Wang, Wenguan , year =. A Survey on. arXiv:2401.03890 , eprint =

  25. [33]

    Text-to-

    Chen, Zilong and Wang, Feng and Liu, Huaping , year =. Text-to-

  26. [34]

    2405.02280 , archiveprefix =

    Chu, Wen-Hsuan and Ke, Lei and Fragkiadaki, Katerina , year =. 2405.02280 , archiveprefix =

  27. [35]

    Vision Transformers Need Registers , booktitle =

    Darcet, Timoth. Vision Transformers Need Registers , booktitle =. 2024 , eprint =

  28. [36]

    Objaverse:

    Deitke, Matt and Schwenk, Dustin and Salvador, Jordi and Weihs, Luca and Michel, Oscar and VanderBilt, Eli and Schmidt, Ludwig and Ehsani, Kiana and Kembhavi, Aniruddha and Farhadi, Ali , year =. Objaverse:. 2212.08051 , archiveprefix =

  29. [37]

    Deng, Yu and Yang, Jiaolong and Xiang, Jianfeng and Tong, Xin , year =

  30. [38]

    and Yan, Xinchen and Zhou, Yin and Guibas, Leonidas and Anguelov, Dragomir , year =

    Deng, Congyue and Jiang, Chiyu Max and Qi, Charles R. and Yan, Xinchen and Zhou, Yin and Guibas, Leonidas and Anguelov, Dragomir , year =

  31. [39]

    Variational

    Deng, Wei and Luo, Weijian and Tan, Yixin and Bilo. Variational. 2024 , eprint =

  32. [40]

    Text-to-

    Ding, Lihe and Dong, Shaocong and Huang, Zhanpeng and Wang, Zibin and Zhang, Yiyuan and Gong, Kaixiong and Xu, Dan and Xue, Tianfan , year =. Text-to-. 2312.04963 , archiveprefix =

  33. [41]

    Dong, Shaocong and Ding, Lihe and Huang, Zhanpeng and Wang, Zibin and Xue, Tianfan and Xu, Dan , year =

  34. [42]

    and Vanhoucke, Vincent , year =

    Downs, Laura and Francis, Anthony and Koenig, Nate and Kinman, Brandon and Hickman, Ryan and Reymann, Krista and McHugh, Thomas B. and Vanhoucke, Vincent , year =. Google. 2204.11918 , archiveprefix =

  35. [43]

    Compositional Visual Generation with Energy Based Models , booktitle =

    Du, Yilun and Li, Shuang and Mordatch, Igor , year =. Compositional Visual Generation with Energy Based Models , booktitle =

  36. [44]

    Progressive

    Dundar, Aysegul and Gao, Jun and Tao, Andrew and Catanzaro, Bryan , year =. Progressive. IEEE Trans. Pattern Anal. Mach. Intell. , volume =

  37. [45]

    and Wu, Jiajun , year =

    Du, Yilun and Zhang, Yinan and Yu, Hong-Xing and Tenenbaum, Joshua B. and Wu, Jiajun , year =. Neural

  38. [46]

    arXiv:2311.10123 , eprint =

    Feng, Lincong and Wang, Muyu and Wang, Maoyu and Xu, Kuo and Liu, Xiaoli , year =. arXiv:2311.10123 , eprint =

  39. [47]

    and Wang, Xiaolong , year =

    Fu, Yang and Liu, Sifei and Kulkarni, Amey and Kautz, Jan and Efros, Alexei A. and Wang, Xiaolong , year =

  40. [48]

    Gao, Chen and Saraf, Ayush and Kopf, Johannes and Huang, Jia-Bin , year =. Dynamic. 2105.06468 , archiveprefix =

  41. [49]

    Frequency-

    Gao, Xiang and Xu, Zhengbo and Zhao, Junhan and Liu, Jiaying , year =. Frequency-. AAAI , volume =

  42. [50]

    arXiv:2403.12365 , eprint =

    Gao, Quankai and Xu, Qiangeng and Cao, Zhe and Mildenhall, Ben and Ma, Wenchao and Chen, Le and Tang, Danhang and Neumann, Ulrich , year =. arXiv:2403.12365 , eprint =

  43. [51]

    Gao, Gege and Liu, Weiyang and Chen, Anpei and Geiger, Andreas and Sch

  44. [52]

    Expressive

    Ge, Songwei and Park, Taesung and Zhu, Jun-Yan and Huang, Jia-Bin , year =. Expressive

  45. [53]

    2023 , eprint =

    Geyer, Michal and. 2023 , eprint =

  46. [54]

    2023 , journal =

    A Spatio-Temporal Network for Video Semantic Segmentation in Surgical Videos , author =. 2023 , journal =. 2306.11052 , archiveprefix =

  47. [55]

    2024 , journal =

    Gu. 2024 , journal =

  48. [56]

    Generating

    Guo, Chuan and Zou, Shihao and Zuo, Xinxin and Wang, Sen and Ji, Wei and Li, Xingyu and Cheng, Li , year =. Generating

  49. [57]

    Point-Bind & Point-

    Guo, Ziyu and Zhang, Renrui and Zhu, Xiangyang and Tang, Yiwen and Ma, Xianzheng and Han, Jiaming and Chen, Kexin and Gao, Peng and Li, Xianzhi and Li, Hongsheng and Heng, Pheng-Ann , year =. Point-Bind & Point-. arXiv:2309.00615 , eprint =

  50. [58]

    arXiv:2405.20283 , eprint =

    Guo, Minghao and Wang, Bohan and He, Kaiming and Matusik, Wojciech , year =. arXiv:2405.20283 , eprint =

  51. [59]

    2023 , howpublished =

    Threestudio , author =. 2023 , howpublished =

  52. [60]

    Hertz, Amir and Aberman, Kfir and. Delta

  53. [61]

    Prompt-to-

    Hertz, Amir and Mokady, Ron and Tenenbaum, Jay and Aberman, Kfir and Pritch, Yael and. Prompt-to-

  54. [62]

    Style Aligned Image Generation via Shared Attention , booktitle =

    Hertz, Amir and Voynov, Andrey and Fruchter, Shlomi and. Style Aligned Image Generation via Shared Attention , booktitle =. 2024 , eprint =

  55. [63]

    He, Yingqing and Yang, Shaoshu and Chen, Haoxin and Cun, Xiaodong and Xia, Menghan and Zhang, Yong and Wang, Xintao and He, Ran and Chen, Qifeng and Shan, Ying , year =

  56. [64]

    He, Yuze and Bai, Yushi and Lin, Matthieu and Zhao, Wang and Hu, Yubin and Sheng, Jenny and Yi, Ran and Li, Juanzi and Liu, Yong-Jin , year =. T\. arXiv:2310.02977 , eprint =

  57. [65]

    He, Yuze and Bai, Yushi and Lin, Matthieu and Sheng, Jenny and Hu, Yubin and Wang, Qi and Wen, Yu-Hui and Liu, Yong-Jin , year =. Text-

  58. [66]

    Classifier-

    Ho, Jonathan and Salimans, Tim , year =. Classifier-

  59. [67]

    Denoising

    Ho, Jonathan and Jain, Ajay and Abbeel, Pieter , year =. Denoising

  60. [68]

    Debiasing

    Hong, Susung and Ahn, Donghoon and Kim, Seungryong , year =. Debiasing

  61. [69]

    Huang, Tianyu and Zhang, Haoze and Zeng, Yihan and Zhang, Zhilu and Li, Hui and Zuo, Wangmeng and Lau, Rynson W H , year =

  62. [70]

    Dreamtime: An Improved Optimization Strategy for Diffusion-Guided 3d Generation , booktitle =

    Huang, Yukun and Wang, Jianan and Shi, Yukai and Tang, Boshi and Qi, Xianbiao and Zhang, Lei , year =. Dreamtime: An Improved Optimization Strategy for Diffusion-Guided 3d Generation , booktitle =

  63. [71]

    Huang, Yukun and Wang, Jianan and Zeng, Ailing and Cao, He and Qi, Xianbiao and Shi, Yukai and Zha, Zheng-Jun and Zhang, Lei , year =

  64. [72]

    Huang, Zehuan and Wen, Hao and Dong, Junting and Wang, Yaohui and Li, Yangguang and Chen, Xinyuan and Cao, Yan-Pei and Liang, Ding and Qiao, Yu and Dai, Bo and Sheng, Lu , year =

  65. [73]

    Huang, Ziqi and Wu, Tianxing and Jiang, Yuming and Chan, Kelvin C. K. and Liu, Ziwei , year =

  66. [74]

    2311.17982 , archiveprefix =

    Huang, Ziqi and He, Yinan and Yu, Jiashuo and Zhang, Fan and Si, Chenyang and Jiang, Yuming and Zhang, Yuanhan and Wu, Tianxing and Jin, Qingyang and Chanpaisit, Nattapol and Wang, Yaohui and Chen, Xinyuan and Wang, Limin and Lin, Dahua and Qiao, Yu and Liu, Ziwei , year =. 23...

  67. [75]

    2403.12002 , archiveprefix =

    Jeong, Hyeonho and Chang, Jinho and Park, Geon Yeong and Ye, Jong Chul , year =. 2403.12002 , archiveprefix =

  68. [76]

    Good Features to Track , booktitle =

  69. [77]

    2311.02848 , archiveprefix =

    Jiang, Yanqin and Zhang, Li and Gao, Jin and Hu, Weimin and Yao, Yao , year =. 2311.02848 , archiveprefix =

  70. [78]

    Efficient-

    Jiang, Yifan and Tang, Hao and Chang, Jen-Hao Rick and Song, Liangchen and Wang, Zhangyang and Cao, Liangliang , year =. Efficient-

  71. [79]

    Jun, Heewoo and Nichol, Alex , year =. Shap-. arXiv:2305.02463 , eprint =

  72. [80]

    ACM Trans

    Kerbl, Bernhard and Kopanas, Georgios and Leimkuehler, Thomas and Drettakis, George , year =. ACM Trans. Graph. , volume =

  73. [81]

    Refining Generative Process with Discriminator Guidance in Score-Based Diffusion Models , booktitle =

    Kim, Dongjun and Kim, Yeongmin and Kwon, Se Jung and Kang, Wanmo and Moon, Il-Chul , year =. Refining Generative Process with Discriminator Guidance in Score-Based Diffusion Models , booktitle =

  74. [82]

    Variational

    Kingma, Diederik P and Salimans, Tim and Poole, Ben and Ho, Jonathan , booktitle =. Variational

  75. [83]

    Dream-in-Style: Text-to-

    Kompanowski, Hubert and Hua, Binh-Son , year =. Dream-in-Style: Text-to-. arXiv:2406.18581 , eprint =

  76. [84]

    , year =

    Krizhevsky, Alex and Sutskever, Ilya and Hinton, Geoffrey E. , year =. Commun. ACM , volume =

  77. [85]

    Multi-Concept Customization of Text-to-Image Diffusion , booktitle =

    Kumari, Nupur and Zhang, Bingliang and Zhang, Richard and Shechtman, Eli and Zhu, Jun-Yan , year =. Multi-Concept Customization of Text-to-Image Diffusion , booktitle =

  78. [86]

    and Kurtek, Sebastian and Bennamoun, Mohammed and Srivastava, Anuj , year =

    Laga, Hamid and Padilla, Marcel and Jermyn, Ian H. and Kurtek, Sebastian and Bennamoun, Mohammed and Srivastava, Anuj , year =. IEEE Trans. Pattern Anal. Mach. Intell. , eprint =

  79. [87]

    Learning

    Lai, Wei-Sheng and Huang, Jia-Bin and Wang, Oliver and Shechtman, Eli and Yumer, Ersin and Yang, Ming-Hsuan , year =. Learning. 1808.00449 , archiveprefix =

  80. [88]

    2311.12050 , archiveprefix =

    Li, Haoran and Ma, Long and Shi, Haolin and Hao, Yanbin and Liao, Yong and Cheng, Lechao and Zhou, Pengyuan , year =. 2311.12050 , archiveprefix =

  81. [89]

    arXiv:2406.13527 , eprint =

    Li, Renjie and Pan, Panwang and Yang, Bangbang and Xu, Dejia and Zhou, Shijie and Zhang, Xuanyang and Li, Zeming and Kadambi, Achuta and Wang, Zhangyang and Tu, Zhengzhong and Fan, Zhiwen , year =. arXiv:2406.13527 , eprint =

  82. [90]

    Advances in

    Li, Xiaoyu and Zhang, Qi and Kang, Di and Cheng, Weihao and Gao, Yiming and Zhang, Jingbo and Liang, Zhihao and Liao, Jing and Cao, Yan-Pei and Shan, Ying , year =. Advances in. arXiv:2401.17807 , eprint =

  83. [91]

    and Zhao, Yao and Wei, Yunchao , year =

    Liang, Hanwen and Yin, Yuyang and Xu, Dejia and Liang, Hanxue and Wang, Zhangyang and Plataniotis, Konstantinos N. and Zhao, Yao and Wei, Yunchao , year =. 2405.16645 , archiveprefix =

  84. [92]

    Liang, Feng and Wu, Bichen and Wang, Jialiang and Yu, Licheng and Li, Kunpeng and Zhao, Yinan and Misra, Ishan and Huang, Jia-Bin and Zhang, Peizhao and Vajda, Peter and Marculescu, Diana , year =

  85. [93]

    Liang, Yixun and Yang, Xin and Lin, Jiantao and Li, Haodong and Xu, Xiaogang and Chen, Yingcong , year =

  86. [94]

    2402.10739 , archiveprefix =

    Liang, Dingkang and Zhou, Xin and Xu, Wei and Zhu, Xingkui and Zou, Zhikang and Ye, Xiaoqing and Tan, Xiao and Bai, Xiang , year =. 2402.10739 , archiveprefix =

  87. [95]

    Reliability-

    Liao, Weihang and. Reliability-. 2023 , journal =

  88. [96]

    Li, Zhiqi and Chen, Yiming and Liu, Peidong , year =

  89. [97]

    2404.03575 , archiveprefix =

    Li, Haoran and Shi, Haolin and Zhang, Wenli and Wu, Wenjun and Liao, Yong and Wang, Lin and Lee, Lik-hang and Zhou, Pengyuan , year =. 2404.03575 , archiveprefix =

  90. [98]

    Generative

    Li, Chenghao and Zhang, Chaoning and Waghwase, Atish and Lee, Lik-Hang and Rameau, Francois and Yang, Yang and Bae, Sung-Ho and Hong, Choong Seon , year =. Generative. arXiv:2305.06131 , eprint =

  91. [99]

    Gnerp: Gaussian-Guided Neural Reconstruc- Tion of Reflective Objects with Noisy Polar- Ization Priors , booktitle =

    Li, Yang and Wu, Ruizheng and Li, Jiyong and Chen, Yingcong , year =. Gnerp: Gaussian-Guided Neural Reconstruc- Tion of Reflective Objects with Noisy Polar- Ization Priors , booktitle =

  92. [100]

    Li, Ming and Zhou, Pan and Liu, Jia-Wei and Keppo, Jussi and Lin, Min and Yan, Shuicheng and Xu, Xiangyu , year =

  93. [101]

    arXiv:2408.03326 , eprint =

    Li, Bo and Zhang, Yuanhan and Guo, Dong and Zhang, Renrui and Li, Feng and Zhang, Hao and Zhang, Kaichen and Zhang, Peiyuan and Li, Yanwei and Liu, Ziwei and Li, Chunyuan , year =. arXiv:2408.03326 , eprint =

  94. [102]

    Li, Tianye and Slavcheva, Mira and Zollhoefer, Michael and Green, Simon and Lassner, Christoph and Kim, Changil and Schmidt, Tanner and Lovegrove, Steven and Goesele, Michael and Newcombe, Richard and Lv, Zhaoyang , year =. Neural. 2103.02597 , archiveprefix =

  95. [103]

    Align Your Gaussians: Text-to-

    Ling, Huan and Kim, Seung Wook and Torralba, Antonio and Fidler, Sanja and Kreis, Karsten , year =. Align Your Gaussians: Text-to-. 2312.13763 , archiveprefix =

  96. [104]

    Gaussian-

    Lin, Youtian and Dai, Zuozhuo and Zhu, Siyu and Yao, Yao , year =. Gaussian-

  97. [105]

    Lin, Chen-Hsuan and Gao, Jun and Tang, Luming and Takikawa, Towaki and Zeng, Xiaohui and Huang, Xun and Kreis, Karsten and Fidler, Sanja and Liu, Ming-Yu and Lin, Tsung-Yi , year =

  98. [106]

    2407.10876 , archiveprefix =

    Li, Chunliang and Han, Wencheng and Yin, Junbo and Zhao, Sanyuan and Shen, Jianbing , year =. 2407.10876 , archiveprefix =

  99. [107]

    Liu, Xi and Zhou, Chaoyi and Huang, Siyu , year =

  100. [108]

    Directional

    Liu, Shengqi and Chen, Zhuo and Gao, Jingnan and Yan, Yichao and Zhu, Wenhan and Lyu, Jiangjing and Yang, Xiaokang , year =. Directional. Computer Graphics Forum , volume =. 2309.14872 , archiveprefix =

  101. [109]

    Liu, Jia-Wei and Cao, Yan-Pei and Wu, Jay Zhangjie and Mao, Weijia and Gu, Yuchao and Zhao, Rui and Keppo, Jussi and Shan, Ying and Shou, Mike Zheng , year =

  102. [110]

    Liu, Zhen and Feng, Yao and Xiu, Yuliang and Liu, Weiyang and Paull, Liam and Black, Michael J. and Sch. Ghost on the

  103. [111]

    2405.12218 , archiveprefix =

    Liu, Tianqi and Wang, Guangcong and Hu, Shoukang and Shen, Liao and Ye, Xinyi and Zang, Yuhang and Cao, Zhiguo and Li, Wei and Liu, Ziwei , year =. 2405.12218 , archiveprefix =

  104. [112]

    Li, Yunxin and Jiang, Shenyuan and Hu, Baotian and Wang, Longyue and Zhong, Wanqi and Luo, Wenhan and Ma, Lin and Zhang, Min , year =. Uni-. IEEE Trans. Pattern Anal. Mach. Intell. , volume =

  105. [113]

    Part123:

    Liu, Anran and Lin, Cheng and Liu, Yuan and Long, Xiaoxiao and Dou, Zhiyang and Guo, Hao-Xiang and Luo, Ping and Wang, Wenping , year =. Part123:. 2405.16888 , archiveprefix =

  106. [114]

    Portrait Diffusion: Training-Free Face Stylization with Chain-of-Painting , shorttitle =

    Liu, Jin and Huang, Huaibo and Jin, Chao and He, Ran , year =. Portrait Diffusion: Training-Free Face Stylization with Chain-of-Painting , shorttitle =. arXiv:2312.02212 , eprint =

  107. [115]

    arXiv.2502.09615 , eprint =

    Liu, Isabella and Xu, Zhan and Yifan, Wang and Tan, Hao and Xu, Zexiang and Wang, Xiaolong and Su, Hao and Shi, Zifan , year =. arXiv.2502.09615 , eprint =

  108. [116]

    Liu, Fangfu and Wu, Diankun and Wei, Yi and Rao, Yongming and Duan, Yueqi , year =

  109. [117]

    Liu, Kunhao and Zhan, Fangneng and Chen, Yiwen and Zhang, Jiahui and Yu, Yingchen and Saddik, Abdulmotaleb El and Lu, Shijian and Xing, Eric , year =

  110. [118]

    Liu, Yuan and Lin, Cheng and Zeng, Zijiao and Long, Xiaoxiao and Liu, Lingjie and Komura, Taku and Wang, Wenping , year =

  111. [119]

    2312.08754 , archiveprefix =

    Liu, Zexiang and Li, Yangguang and Lin, Youtian and Yu, Xin and Peng, Sida and Cao, Yan-Pei and Qi, Xiaojuan and Huang, Xiaoshui and Liang, Ding and Ouyang, Wanli , year =. 2312.08754 , archiveprefix =

  112. [120]

    Liu, Haotian and Li, Chunyuan and Wu, Qingyang and Lee, Yong Jae , year =. Visual

  113. [121]

    2408.05492 , archiveprefix =

    Liu, Jin and Huang, Huaibo and Cao, Jie and He, Ran , year =. 2408.05492 , archiveprefix =

  114. [122]

    Zero-1-to-3:

    Liu, Ruoshi and Wu, Rundi and Hoorick, Basile Van and Tokmakov, Pavel and Zakharov, Sergey and Vondrick, Carl , year =. Zero-1-to-3:

  115. [123]

    2310.15008 , archiveprefix =

    Long, Xiaoxiao and Guo, Yuan-Chen and Lin, Cheng and Liu, Yuan and Dou, Zhiyang and Liu, Lingjie and Ma, Yuexin and Zhang, Song-Hai and Habermann, Marc and Theobalt, Christian and Wang, Wenping , year =. 2310.15008 , archiveprefix =

  116. [124]

    Lorraine, Jonathan and Xie, Kevin and Zeng, Xiaohui and Lin, Chen-Hsuan and Takikawa, Towaki and Sharp, Nicholas and Lin, Tsung-Yi and Liu, Ming-Yu and Fidler, Sanja and Lucas, James , year =

  117. [125]

    Lucas, Bruce D and Kanade, Takeo , year =. An

  118. [126]

    Luiten, Jonathon and Kopanas, Georgios and Leibe, Bastian and Ramanan, Deva , year =. Dynamic

  119. [127]

    Scalable

    Luo, Tiange and Rockwell, Chris and Lee, Honglak and Johnson, Justin , year =. Scalable

  120. [128]

    arXiv.org , volume =

    Luo, Chaofan and Di, Donglin and Yang, Xun and Ma, Yongjia and Xue, Zhou and Wei, Chen and Liu, Yebin , year =. arXiv.org , volume =. 2407.02034 , archiveprefix =

  121. [129]

    Scaffold-

    Lu, Tao and Yu, Mulin and Xu, Linning and Xiangli, Yuanbo and Wang, Limin and Lin, Dahua and Dai, Bo , year =. Scaffold-. 2312.00109 , archiveprefix =

  122. [130]

    Conditional

    della Maggiora, Gabriel and Croquevielle, Luis Alberto and Deshpande, Nikita and Horsley, Harry and Heinis, Thomas and Yakimovich, Artur , year =. Conditional. 2312.02246 , archiveprefix =

  123. [131]

    arXiv:2401.07727 , eprint =

    Mercier, Antoine and Nakhli, Ramin and Reddy, Mahesh and Yasarla, Rajeev and Cai, Hong and Porikli, Fatih and Berger, Guillaume , year =. arXiv:2401.07727 , eprint =

  124. [132]

    Michel, Oscar and Bhattad, Anand and VanderBilt, Eli and Krishna, Ranjay and Kembhavi, Aniruddha and Gupta, Tanmay , year =

  125. [133]

    and Tancik, Matthew and Barron, Jonathan T

    Mildenhall, Ben and Srinivasan, Pratul P. and Tancik, Matthew and Barron, Jonathan T. and Ramamoorthi, Ravi and Ng, Ren , year =. Commun. ACM , lccn =

  126. [134]

    Null-Text

    Mokady, Ron and Hertz, Amir and Aberman, Kfir and Pritch, Yael and. Null-Text

  127. [135]

    Contrastive Denoising Score for Text-Guided Latent Diffusion Image Editing , booktitle =

    Nam, Hyelin and Kwon, Gihyun and Park, Geon Yeong and Ye, Jong Chul , year =. Contrastive Denoising Score for Text-Guided Latent Diffusion Image Editing , booktitle =. 2311.18608 , archiveprefix =

  128. [136]

    Nam, Jisu and Kim, Heesu and Lee, DongJae and Jin, Siyoon and Kim, Seungryong and Chang, Seunggyu , year =

  129. [137]

    Nichol, Alex and Jun, Heewoo and Dhariwal, Prafulla and Mishkin, Pamela and Chen, Mark , year =. Point-. arXiv:2212.08751 , eprint =

  130. [138]

    Tong, Haoyang and Wang, Hongbo and Liu, Jin and Wang, Qi and Cao, Jie and He, Ran , year =

  131. [139]

    Liu, Xinyue and Liu, Jin and Wang, Hongbo and He, Ran and Huang, Huaibo , year =. Think-

  132. [140]

    Sun, Jiayang and Wang, Pin and Wang, Hongbo and Liu, Xinyue and Huang, Huaibo and He, Ran , year =. Towards

  133. [141]

    Machine Intelligence Research , author =

    Marmot:. Machine Intelligence Research , author =. 2026 , eprint =

  134. [142]

    2025 , eprint =

    IJCV , author =. 2025 , eprint =

  135. [143]

    IJCV , author =

    Hyper-. IJCV , author =. 2025 , eprint =

  136. [144]

    IJCV , author =

    Diffusion. IJCV , author =. 2025 , eprint =

  137. [145]

    Coloring the

    Wang, Hongbo and Huang, Huaibo and Wang, Pin and Hao, Jinhua and Zhou, Chao and He, Ran , year =. Coloring the. 2605.23264 , archiveprefix =

  138. [146]

    Structured

    Xiang, Jianfeng and Lv, Zelong and Xu, Sicheng and Deng, Yu and Wang, Ruicheng and Zhang, Bowen and Chen, Dong and Tong, Xin and Yang, Jiaolong , year =. Structured. 2412.01506 , archiveprefix =

  139. [147]

    and Holynski, Aleksander , year =

    Wu, Rundi and Gao, Ruiqi and Poole, Ben and Trevithick, Alex and Zheng, Changxi and Barron, Jonathan T. and Holynski, Aleksander , year =. 2411.18613 , archiveprefix =

  140. [148]

    2512.23568 , archiveprefix =

    Jiao, Siyu and Lin, Yiheng and Zhong, Yujie and She, Qi and Zhou, Wei and Lan, Xiaohan and Huang, Zilong and Yu, Fei and Yu, Yingchen and Zhao, Yunqing and Zhao, Yao and Wei, Yunchao , year =. 2512.23568 , archiveprefix =

  141. [149]

    Thinking-While-

    Guo, Ziyu and Zhang, Renrui and Li, Hongyu and Zhang, Manyuan and Chen, Xinyan and Wang, Sifan and Feng, Yan and Pei, Peng and Heng, Pheng-Ann , year =. Thinking-While-. 2511.16671 , archiveprefix =

  142. [150]

    2026 , eprint =

    arXiv:2605.01743 , author =. 2026 , eprint =

  143. [151]

    2024 , journal =

    Oquab, Maxime and Darcet, Timoth. 2024 , journal =. 2304.07193 , archiveprefix =

  144. [152]

    Ouyang, Hao and Wang, Qiuyu and Xiao, Yuxi and Bai, Qingyan and Zhang, Juntao and Zheng, Kecheng and Zhou, Xiaowei and Chen, Qifeng and Chen, Qifeng , year =

  145. [153]

    arXiv:2312.09242 , eprint =

    Ouyang, Hao and Heal, Kathryn and Lombardi, Stephen and Sun, Tiancheng , year =. arXiv:2312.09242 , eprint =

  146. [154]

    Diffusion Handles: Enabling

    Pandey, Karran and Guerrero, Paul and Gadelha, Matheus and. Diffusion Handles: Enabling. 2023 , eprint =

  147. [155]

    and Zhang, Richard and Zhu, Jun-Yan , year =

    Park, Taesung and Efros, Alexei A. and Zhang, Richard and Zhu, Jun-Yan , year =. Contrastive Learning for Unpaired Image-to-Image Translation , booktitle =. 2007.15651 , archiveprefix =

  148. [156]

    and Bouaziz, Sofien and Goldman, Dan B

    Park, Keunhong and Sinha, Utkarsh and Barron, Jonathan T. and Bouaziz, Sofien and Goldman, Dan B. and Seitz, Steven M. and. Nerfies:. 2020 , eprint =

  149. [157]

    Zero-Shot Image-to-Image Translation , booktitle =

    Parmar, Gaurav and Singh, Krishna Kumar and Zhang, Richard and Li, Yijun and Lu, Jingwan and Zhu, Jun-Yan , year =. Zero-Shot Image-to-Image Translation , booktitle =. 2302.03027 , archiveprefix =

  150. [158]

    and Mildenhall, Ben , year =

    Poole, Ben and Jain, Ajay and Barron, Jonathan T. and Mildenhall, Ben , year =

  151. [159]

    2024 , journal =

    State of the Art on Diffusion Models for Visual Computing , author =. 2024 , journal =. 2310.07204 , archiveprefix =

  152. [160]

    Pumarola, Albert and Corona, Enric and. D-. 2020 , eprint =

  153. [161]

    Magic123:

    Qian, Guocheng and Mai, Jinjie and Hamdi, Abdullah and Ren, Jian and Siarohin, Aliaksandr and Li, Bing and Lee, Hsin-Ying and Skorokhodov, Ivan and Wonka, Peter and Tulyakov, Sergey and Ghanem, Bernard , year =. Magic123:

  154. [162]

    Qin, Minghan and Li, Wanhua and Zhou, Jiawei and Wang, Haoqian and Pfister, Hanspeter , year =

  155. [163]

    IEEE Trans

    Qiu, Shi and Anwar, Saeed and Barnes, Nick , year =. IEEE Trans. Pattern Anal. Mach. Intell. , volume =

  156. [164]

    2311.16918 , archiveprefix =

    Qiu, Lingteng and Chen, Guanying and Gu, Xiaodong and Zuo, Qi and Xu, Mutian and Wu, Yushuang and Yuan, Weihao and Dong, Zilong and Bo, Liefeng and Han, Xiaoguang , year =. 2311.16918 , archiveprefix =

  157. [165]

    Learning

    Radford, Alec and Kim, Jong Wook and Hallacy, Chris and Ramesh, Aditya and Goh, Gabriel and Agarwal, Sandhini and Sastry, Girish and Askell, Amanda and Mishkin, Pamela and Clark, Jack and Krueger, Gretchen and Sutskever, Ilya , year =. Learning

  158. [166]

    Raj, Amit and Kaza, Srinivas and Poole, Ben and Niemeyer, Michael and Ruiz, Nataniel and Mildenhall, Ben and Zada, Shiran and Aberman, Kfir and Rubinstein, Michael and Barron, Jonathan and Li, Yuanzhen and Jampani, Varun , year =

  159. [167]

    arXiv:2312.17142 , volume =

    Ren, Jiawei and Pan, Liang and Tang, Jiaxiang and Zhang, Chi and Cao, Ang and Zeng, Gang and Liu, Ziwei , year =. arXiv:2312.17142 , volume =. 2312.17142 , archiveprefix =

  160. [168]

    2023 , eprint =

    Richardson, Elad and Metzer, Gal and Alaluf, Yuval and Giryes, Raja and. 2023 , eprint =

  161. [169]

    High-Resolution Image Synthesis with Latent Diffusion Models , booktitle =

    Rombach, Robin and Blattmann, Andreas and Lorenz, Dominik and Esser, Patrick and Ommer, Bj. High-Resolution Image Synthesis with Latent Diffusion Models , booktitle =

  162. [170]

    Ronneberger, Olaf and Fischer, Philipp and Brox, Thomas , year =. U-. 1505.04597 , archiveprefix =

  163. [171]

    arXiv:2405.17401 , eprint =

    Rout, Litu and Chen, Yujia and Ruiz, Nataniel and Kumar, Abhishek and Caramanis, Constantine and Shakkottai, Sanjay and Chu, Wen-Sheng , year =. arXiv:2405.17401 , eprint =

  164. [172]

    Ruiz, Nataniel and Li, Yuanzhen and Jampani, Varun and Pritch, Yael and Rubinstein, Michael and Aberman, Kfir , year =

  165. [173]

    Photorealistic

    Saharia, Chitwan and Chan, William and Saxena, Saurabh and Li, Lala and Whang, Jay and Denton, Emily and Ghasemipour, Seyed Kamyar Seyed and Ayan, Burcu Karagol and Mahdavi, S Sara and. Photorealistic

  166. [174]

    2502.09278 , archiveprefix =

    2025 , journal =. 2502.09278 , archiveprefix =

  167. [175]

    Sella, Etai and Fiebelman, Gal and Hedman, Peter and. Vox-

  168. [176]

    arXiv:2304.02827 , eprint =

    Seo, Hoigi and Kim, Hayeon and Kim, Gwanghyun and Chun, Se Young , year =. arXiv:2304.02827 , eprint =

  169. [177]

    Seo, Junyoung and Jang, Wooseok and Kwak, Min-Seop and Ko, Jaehoon and Kim, Hyeonsu and Kim, Junho and Kim, Jin-Hwa and Lee, Jiyoung and Kim, Seungryong , year =. Let

  170. [178]

    2024 , journal =

    Learning Temporally Consistent Video Depth from Video Diffusion Priors , author =. 2024 , journal =. 2406.01493 , archiveprefix =

  171. [179]

    Shi, Yujun and Xue, Chuhui and Liew, Jun Hao and Pan, Jiachun and Yan, Hanshu and Zhang, Wenqing and Tan, Vincent Y. F. and Bai, Song , year =. 2306.14435 , archiveprefix =

  172. [180]

    Shi, Yichun and Wang, Peng and Ye, Jianglong and Long, Mai and Li, Kejie and Yang, Xiao , year =

  173. [181]

    Zero123++: A

    Shi, Ruoxi and Chen, Hansheng and Zhang, Zhuoyang and Liu, Minghua and Xu, Chao and Wei, Xinyue and Chen, Linghao and Zeng, Chong and Su, Hao , year =. Zero123++: A. arXiv:2310.15110 , eprint =

  174. [182]

    Si, Chenyang and Huang, Ziqi and Jiang, Yuming and Liu, Ziwei , year =

  175. [183]

    Simonyan, Karen and Zisserman, Andrew , year =. Very

  176. [184]

    Make-a-Video: Text-to-Video Generation without Text-Video Data , shorttitle =

    Singer, Uriel and Polyak, Adam and Hayes, Thomas and Yin, Xi and An, Jie and Zhang, Songyang and Hu, Qiyuan and Yang, Harry and Ashual, Oron and Gafni, Oran and Parikh, Devi and Gupta, Sonal and Taigman, Yaniv , year =. Make-a-Video: Text-to-Video Generation without Text-Video...

  177. [185]

    Text-to-

    Singer, Uriel and Sheynin, Shelly and Polyak, Adam and Ashual, Oron and Makarov, Iurii and Kokkinos, Filippos and Goyal, Naman and Vedaldi, Andrea and Parikh, Devi and Johnson, Justin and Taigman, Yaniv , year =. Text-to-. 2301.11280 , archiveprefix =

  178. [186]

    Denoising

    Song, Jiaming and Meng, Chenlin and Ermon, Stefano , year =. Denoising

  179. [187]

    arXiv:2411.04928 , eprint =

    Sun, Wenqiang and Chen, Shuo and Liu, Fangfu and Chen, Zilong and Duan, Yueqi and Zhang, Jun and Wang, Yikai , year =. arXiv:2411.04928 , eprint =

  180. [188]

    Splatter Image: Ultra-Fast Single-View

  181. [189]

    High-Resolution Image Reconstruction with Latent Diffusion Models from Human Brain Activity , booktitle =

    Takagi, Yu and Nishimoto, Shinji , year =. High-Resolution Image Reconstruction with Latent Diffusion Models from Human Brain Activity , booktitle =

  182. [190]

    Tang, Jiaxiang and Ren, Jiawei and Zhou, Hang and Liu, Ziwei and Zeng, Gang , year =

  183. [191]

    Emergent Correspondence from Image Diffusion , booktitle =

    Tang, Luming and Jia, Menglin and Wang, Qianqian and Phoo, Cheng Perng and Hariharan, Bharath , year =. Emergent Correspondence from Image Diffusion , booktitle =. 2306.03881 , archiveprefix =

  184. [192]

    2402.05054 , archiveprefix =

    Tang, Jiaxiang and Chen, Zhaoxi and Chen, Xiaokang and Wang, Tengfei and Zeng, Gang and Liu, Ziwei , year =. 2402.05054 , archiveprefix =

  185. [193]

    Tang, Junshu and Wang, Tengfei and Zhang, Bo and Zhang, Ting and Yi, Ran and Ma, Lizhuang and Chen, Dong , year =. Make-

  186. [194]

    Tewel, Yoad and Gal, Rinon and Chechik, Gal and Atzmon, Yuval , year =. Key-

  187. [195]

    arXiv:2403.02151 , eprint =

    Tochilkin, Dmitry and Pankratz, David and Liu, Zexiang and Huang, Zixuan and Letts, Adam and Li, Yangguang and Liang, Ding and Laforte, Christian and Jampani, Varun and Cao, Yan-Pei , year =. arXiv:2403.02151 , eprint =

  188. [196]

    Plug-and-

    Tumanyan, Narek and Geyer, Michal and Bagon, Shai and Dekel, Tali , year =. Plug-and-

  189. [197]

    arXiv:2311.17907 , eprint =

    Vilesov, Alexander and Chari, Pradyumna and Kadambi, Achuta , year =. arXiv:2311.17907 , eprint =

  190. [198]

    Voleti, Vikram and Yao, Chun-Han and Boss, Mark and Letts, Adam and Pankratz, David and Tochilkin, Dmitry and Laforte, Christian and Rombach, Robin and Jampani, Varun , year =

  191. [199]

    2406.08850 , archiveprefix =

    Wang, Jiangshan and Ma, Yue and Guo, Jiayi and Xiao, Yicheng and Huang, Gao and Li, Xiu , year =. 2406.08850 , archiveprefix =

  192. [200]

    Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image Editing , booktitle =

    Wang, Kai and Yang, Fei and Yang, Shiqi and Butt, Muhammad Atif and. Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image Editing , booktitle =

  193. [201]

    2407.05600 , archiveprefix =

    Wang, Zhenyu and Li, Aoxue and Li, Zhenguo and Liu, Xihui , year =. 2407.05600 , archiveprefix =

  194. [202]

    Wang, Hongbo and Cao, Jie and Liu, Jin and Zhou, Xiaoqiang and Huang, Huaibo and He, Ran , year =

  195. [203]

    arXiv:2312.02201 , eprint =

    Wang, Peng and Shi, Yichun , year =. arXiv:2312.02201 , eprint =

  196. [204]

    Wang, Hengyi and Wang, Jingwen and Agapito, Lourdes , year =

  197. [205]

    Wang, Zhengyi and Lu, Cheng and Wang, Yikai and Bao, Fan and Li, Chongxuan and Su, Hang and Zhu, Jun , year =

  198. [206]

    and Shakhnarovich, Greg , year =

    Wang, Haochen and Du, Xiaodan and Li, Jiahao and Yeh, Raymond A. and Shakhnarovich, Greg , year =. Score

  199. [207]

    2024 , journal =

    Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation , author =. 2024 , journal =. 2305.10874 , archiveprefix =

  200. [208]

    Wang, Yikai and Liu, Guangce and Wang, Xinzhou and Chen, Zilong and Li, Jiafang and Liang, Xin and Sun, Fuchun and Zhu, Jun , year =

  201. [209]

    2310.08528 , archiveprefix =

    Wu, Guanjun and Yi, Taoran and Fang, Jiemin and Xie, Lingxi and Zhang, Xiaopeng and Wei, Wei and Liu, Wenyu and Tian, Qi and Wang, Xinggang , year =. 2310.08528 , archiveprefix =

  202. [210]

    Wu, Zike and Zhou, Pan and Yi, Xuanyu and Yuan, Xiaoding and Zhang, Hanwang , year =

  203. [211]

    2405.14832 , archiveprefix =

    Wu, Shuang and Lin, Youtian and Zhang, Feihu and Zeng, Yifei and Xu, Jingxi and Torr, Philip and Cao, Xun and Yao, Yao , year =. 2405.14832 , archiveprefix =

  204. [212]

    Wu, Zongwei and Zheng, Jilai and Ren, Xiangxuan and Vasluianu, Florin-Alexandru and Ma, Chao and Paudel, Danda Pani and Van Gool, Luc and Timofte, Radu , year =. Single-

  205. [213]

    2405.20343 , archiveprefix =

    Wu, Kailu and Liu, Fangfu and Cai, Zhihan and Yan, Runjie and Wang, Hanyang and Hu, Yating and Duan, Yueqi and Ma, Kaisheng , year =. 2405.20343 , archiveprefix =

  206. [214]

    Xiang, Jianfeng and Yang, Jiaolong and Huang, Binbin and Tong, Xin , year =

  207. [215]

    arXiv:2310.04561 , eprint =

    Xie, Tianhao and Belilovsky, Eugene and Mudur, Sudhir and Popa, Tiberiu , year =. arXiv:2310.04561 , eprint =

  208. [216]

    2311.12198 , archiveprefix =

    Xie, Tianyi and Zong, Zeshun and Qiu, Yuxing and Li, Xuan and Feng, Yutao and Yang, Yin and Jiang, Chenfanfu , year =. 2311.12198 , archiveprefix =

  209. [217]

    arXiv:2407.17470 , eprint =

    Xie, Yiming and Yao, Chun-Han and Voleti, Vikram and Jiang, Huaizu and Jampani, Varun , year =. arXiv:2407.17470 , eprint =

  210. [218]

    2405.17933 , archiveprefix =

    Xing, Jinbo and Liu, Hanyuan and Xia, Menghan and Zhang, Yong and Wang, Xintao and Shan, Ying and Wong, Tien-Tsin , year =. 2405.17933 , archiveprefix =

  211. [219]

    arXiv:2401.04099 , eprint =

    Xu, Dejia and Yuan, Ye and Mardani, Morteza and Liu, Sifei and Song, Jiaming and Wang, Zhangyang and Vahdat, Arash , year =. arXiv:2401.04099 , eprint =

  212. [220]

    and Hu, Hezhen and Liang, Hanxue and Plataniotis, Konstantinos N

    Xu, Dejia and Liang, Hanwen and Bhatt, Neel P. and Hu, Hezhen and Liang, Hanxue and Plataniotis, Konstantinos N. and Wang, Zhangyang , year =. arXiv:2403.16993 , eprint =

  213. [221]

    2212.14704 , archiveprefix =

    Xu, Jiale and Wang, Xintao and Cheng, Weihao and Cao, Yan-Pei and Shan, Ying and Qie, Xiaohu and Gao, Shenghua , year =. 2212.14704 , archiveprefix =

  214. [222]

    Xu, Yunqiu and Zhu, Linchao and Yang, Yi , year =

  215. [223]

    2403.14621 , archiveprefix =

    Xu, Yinghao and Shi, Zifan and Yifan, Wang and Chen, Hansheng and Yang, Ceyuan and Peng, Sida and Shen, Yujun and Wetzstein, Gordon , year =. 2403.14621 , archiveprefix =

  216. [224]

    arXiv:2404.07191 , eprint =

    Xu, Jiale and Cheng, Weihao and Gao, Yiming and Wang, Xintao and Gao, Shenghua and Shan, Ying , year =. arXiv:2404.07191 , eprint =

  217. [225]

    2023 , journal =

    Inversion-Free Image Editing with Natural Language , author =. 2023 , journal =. 2312.04965 , archiveprefix =

  218. [226]

    Xu, Zhongcong and Zhang, Jianfeng and Liew, Jun Hao and Yan, Hanshu and Liu, Jia-Wei and Zhang, Chenxu and Feng, Jiashi and Shou, Mike Zheng , year =

  219. [227]

    2308.16911 , archiveprefix =

    Xu, Runsen and Wang, Xiaolong and Wang, Tai and Chen, Yilun and Pang, Jiangmiao and Lin, Dahua , year =. 2308.16911 , archiveprefix =

  220. [228]

    2407.16260 , archiveprefix =

    Yan, Zizheng and Zhou, Jiapeng and Meng, Fanpeng and Wu, Yushuang and Qiu, Lingteng and Ye, Zisheng and Cui, Shuguang and Chen, Guanying and Han, Xiaoguang , year =. 2407.16260 , archiveprefix =

  221. [229]

    Yang, Jiayu and Cheng, Ziang and Duan, Yunfei and Ji, Pan and Li, Hongdong , year =

  222. [230]

    Yang, Shuai and Zhou, Yifan and Liu, Ziwei and Loy, Chen Change , year =. Fresco:

  223. [231]

    ACM Trans

    Yang, Chen and Li, Sikuang and Fang, Jiemin and Liang, Ruofan and Xie, Lingxi and Zhang, Xiaopeng and Shen, Wei and Tian, Qi , year =. ACM Trans. Graph. , eprint =

  224. [232]

    Mastering

    Yang, Ling and Yu, Zhaochen and Meng, Chenlin and Xu, Minkai and Ermon, Stefano and Cui, Bin , year =. Mastering. 2401.11708 , archiveprefix =

  225. [233]

    Not All Frame Features Are Equal: Video-to-

    Yang, Liying and Liu, Chen and Zhu, Zhenwei and Liu, Ajian and Ma, Hui and Nong, Jian and Liang, Yanyan , year =. Not All Frame Features Are Equal: Video-to-. arXiv:2502.08377 , eprint =

  226. [234]

    2311.11700 , archiveprefix =

    Yan, Chi and Qu, Delin and Xu, Dan and Zhao, Bin and Wang, Zhigang and Wang, Dong and Li, Xuelong , year =. 2311.11700 , archiveprefix =

  227. [235]

    arXiv:2312.12634 , eprint =

    Yazdian, Payam Jome and Liu, Eric and Lagasse, Rachel and Mohammadi, Hamid and Cheng, Li and Lim, Angelica , year =. arXiv:2312.12634 , eprint =

  228. [236]

    Consistent-1-to-3:

    Ye, Jianglong and Wang, Peng and Li, Kejie and Shi, Yichun and Wang, Heng , year =. Consistent-1-to-3:

  229. [237]

    Gaussian Grouping: Segment and Edit Anything in

    Ye, Mingqiao and Danelljan, Martin and Yu, Fisher and Ke, Lei , year =. Gaussian Grouping: Segment and Edit Anything in. 2312.00732 , archiveprefix =

  230. [238]

    Yi, Taoran and Fang, Jiemin and Wang, Junjie and Wu, Guanjun and Xie, Lingxi and Zhang, Xiaopeng and Liu, Wenyu and Tian, Qi and Wang, Xinggang , year =

  231. [239]

    arXiv:2407.12684 , volume =

    Yuan, Yu-Jie and Kobbelt, Leif and Liu, Jiwen and Zhang, Yuan and Wan, Pengfei and Lai, Yu-Kun and Gao, Lin , year =. arXiv:2407.12684 , volume =. 2407.12684 , archiveprefix =

  232. [240]

    Mip-Splatting: Alias-Free

    Yu, Zehao and Chen, Anpei and Huang, Binbin and Sattler, Torsten and Geiger, Andreas , year =. Mip-Splatting: Alias-Free

  233. [241]

    arXiv:2406.15811 , eprint =

    Yu, Qiao and Li, Xianzhi and Tang, Yuan and Xu, Jinfeng and Hu, Long and Hao, Yixue and Chen, Min , year =. arXiv:2406.15811 , eprint =

  234. [242]

    Text-to-

    Yu, Xin and Guo, Yuan-Chen and Li, Yangguang and Liang, Ding and Zhang, Song-Hai and Qi, Xiaojuan , year =. Text-to-. ICLR , volume =. 2310.19415 , archiveprefix =

  235. [243]

    Zeng, Yan and Wei, Guoqiang and Zheng, Jiani and Zou, Jiaxin and Wei, Yang and Zhang, Yuchen and Li, Hang , year =. Make. 2311.10982 , archiveprefix =

  236. [244]

    2403.14939 , archiveprefix =

    Zeng, Yifei and Jiang, Yanqin and Zhu, Siyu and Lu, Yuanxun and Lin, Youtian and Zhu, Hao and Hu, Weiming and Cao, Xun and Yao, Yao , year =. 2403.14939 , archiveprefix =

  237. [245]

    arXiv:2403.02217 , eprint =

    Zhang, Yudi and Xu, Qi and Zhang, Lei , year =. arXiv:2403.02217 , eprint =

  238. [246]

    and Zheng, Changxi and Snavely, Noah and Wu, Jiajun and Freeman, William T

    Zhang, Tianyuan and Yu, Hong-Xing and Wu, Rundi and Feng, Brandon Y. and Zheng, Changxi and Snavely, Noah and Wu, Jiajun and Freeman, William T. , year =. 2404.13026 , archiveprefix =

  239. [247]

    Zhang, Zicheng and Wu, Haoning and Zhang, Erli and Zhai, Guangtao and Lin, Weisi , year =. Q-. IEEE Trans. Pattern Anal. Mach. Intell. , volume =

  240. [248]

    Real-World Image Variation by Aligning Diffusion Inversion Chain , booktitle =

    Zhang, Yuechen and Xing, Jinbo and Lo, Eric and Jia, Jiaya , year =. Real-World Image Variation by Aligning Diffusion Inversion Chain , booktitle =

  241. [249]

    arXiv:2404.05220 , eprint =

    Zhang, Dingxi and Yuan, Yu-Jie and Chen, Zhuoxun and Zhang, Fang-Lue and He, Zhenliang and Shan, Shiguang and Gao, Lin , year =. arXiv:2404.05220 , eprint =

  242. [250]

    arXiv:2412.16919 , eprint =

    Zhang, Xuying and Liu, Yutong and Li, Yangguang and Zhang, Renrui and Liu, Yufei and Wang, Kai and Ouyang, Wanli and Xiong, Zhiwei and Gao, Peng and Hou, Qibin and Cheng, Ming-Ming , year =. arXiv:2412.16919 , eprint =

  243. [251]

    IEEE Trans

    Zhang, Jingbo and Li, Xiaoyu and Wan, Ziyu and Wang, Can and Liao, Jing , year =. IEEE Trans. Visual. Comput. Graphics , volume =

  244. [252]

    Transparent

    Zhang, Lvmin and Agrawala, Maneesh , year =. Transparent

  245. [253]

    and Shechtman, Eli and Wang, Oliver , year =

    Zhang, Richard and Isola, Phillip and Efros, Alexei A. and Shechtman, Eli and Wang, Oliver , year =. The

  246. [254]

    Zhao, Minda and Zhao, Chaoyi and Liang, Xinyue and Li, Lincheng and Zhao, Zeng and Hu, Zhipeng and Fan, Changjie and Yu, Xin , year =

  247. [255]

    Zheng, Chenxi and Lin, Yihong and Liu, Bangzhen and Xu, Xuemiao and Nie, Yongwei and He, Shengfeng , year =

  248. [256]

    Zheng, Yufeng and Li, Xueting and Nagano, Koki and Liu, Sifei and Hilliges, Otmar and De Mello, Shalini , year =. A

  249. [257]

    Zhou, Linqi and Shih, Andy and Meng, Chenlin and Ermon, Stefano , year =

  250. [258]

    2402.07207 , archiveprefix =

    Zhou, Xiaoyu and Ran, Xingjian and Xiong, Yajiao and He, Jinlin and Lin, Zhiwei and Wang, Yongtao and Sun, Deqing and Yang, Ming-Hsuan , year =. 2402.07207 , archiveprefix =

  251. [259]

    Upscale-

    Zhou, Shangchen and Yang, Peiqing and Wang, Jianyi and Luo, Yihang and Loy, Chen Change , year =. Upscale-

  252. [260]

    2004 , journal =

    Image Quality Assessment: From Error Visibility to Structural Similarity , shorttitle =. 2004 , journal =

  253. [261]

    ACM Trans

    Zhuang, Jingyu and Kang, Di and Cao, Yan-Pei and Li, Guanbin and Lin, Liang and Shan, Ying , year =. ACM Trans. Graph. , eprint =

  254. [262]

    arXiv:2412.15278 , volume =

    Zhu, Xingyu and Luo, Xiapu and Wei, Xuetao , year =. arXiv:2412.15278 , volume =. 2412.15278 , archiveprefix =

  255. [263]

    2305.18766 , archiveprefix =

    Zhu, Junzhe and Zhuang, Peiye and Koyejo, Sanmi , year =. 2305.18766 , archiveprefix =

  256. [264]

    Triplane

    Zou, Zi--Xin and Yu, Zhipeng and Guo, Yuan--Chen and Li, Yangguang and Liang, Ding and Cao, Yan--Pei and Zhang, Song--Hai , year =. Triplane

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.