REVIEW 4 major objections 5 minor 54 references
OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read This paper claims that a ruined film's picture and soundtrack should be restored together in one generative pass, and that the resulting joint model beats every prior single-modality method on all measured visual, audio, and sync metrics.
desk verdict A genuinely novel joint AV restoration system with strong visual results, but the headline audio claim is contradicted by the paper's own appendix and needs a major rewrite. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the multimodal diffusion transformer itself, a 22-billion-parameter network that generates video and audio natively in one shared token stream with cross-modal attention gates. OmniVR keeps the architecture and changes only what enters it: the low-quality video and audio are encoded by frozen VAEs into condition tokens, perturbed with noise, and channel-concatenated with the noisy target tokens through a projection whose new columns start at zero, so the model begins as a functioning text-to-audio-video generator and gradually learns to treat the degraded pair as evidence. The three supporting mechanisms are the joint audio-video degradation pipeline that supplies co-degraded training pairs with per-sample severity scalars, the prompt-annealing schedule that ends in a fixed restoration prompt requiring no captions at inference, and first-frame image-to-video anchoring with degradation-aware loss reweighting and a multi-resolution STFT waveform loss.
What would settle it
Take any real archival film for which a clean reference survives—a surviving color print, the original negative, or a professional re-master—and compare OmniVR with the best single-modality restorers using full-reference metrics (LPIPS and FVD for video; PESQ, STOI, and SI-SDR for audio). The central claim predicts the joint model wins or ties on all of them; the paper's own controlled-track appendix already reports it below an audio-only enhancer on all three waveform-alignment metrics, so if that reversal reproduces on real footage, the joint-superiority claim would need to be restricted to perceptual, no-reference measures.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that the generative prior of a large joint audio-video model can be repurposed into a restoration engine without rebuilding the architecture: the degraded pair is encoded by frozen VAEs into tokens, channel-concatenated with the noisy clean-target tokens, and jointly denoised by the same bidirectional cross-modal attention that originally generated video from text. Three adaptations make this work: a joint degradation pipeline that produces realistically co-degraded training pairs, a prompt-annealing schedule that shifts the text condition from descriptive captions to a single fixed restoration prompt while the condition-projection weights grow from zero, and first-frame anchoring that chains 121-frame windows with a waveform-domain supervision loss. The paper's evidence is OmniVRBench, 200 real historical clips evaluated on visual quality, audio quality, temporal consistency, and audio-visual synchrony, where OmniVR reports leading scores on every metric, the best audio quality, natural colorization, and the best lip-sync among methods that touch the audio at all, including over cascaded video-plus-audio baselines.
Load-bearing premise
The load-bearing premise is that the synthetic joint degradation pipeline, which produces every training pair and every controlled evaluation clip, is a faithful enough proxy for real archival damage that what the model learns on it transfers to authentic historical footage, which is evaluated only with no-reference metrics and has no clean reference at all.
Editorial extensions
If this is right
- Restoration of a historical film becomes a single-model operation: one inference call repairs blur, noise, flicker, and grayscale in the frames while removing hiss, clipping, dropout, and bandwidth loss from the soundtrack, with no per-clip captions required.
- Joint generation is what preserves audio-visual alignment: cascaded combinations of the best separate video and audio restorers match audio-only methods on sound quality but score far worse on lip-sync, showing that alignment cannot be re-assembled later from independent tools.
- Very long films are handled at bounded memory by chaining 121-frame windows, each seeded by the previous window's restored boundary frame, with the very first window restored from the degraded input alone.
- Generative restoration can score higher than the clean reference on no-reference estimators, so the paper reads the full-reference metrics in its appendix as the primary signal for faithfulness rather than as a claim that the output objectively beats the original footage.
Reading between the lines
- A testable extension implied by the paper's cross-modal routing observations: measure how much of an acoustically ambiguous track, such as speech under heavy dropout, the joint model can recover from lip motion alone, compared with an audio-only restorer fed the same degraded input.
- The paper's own appendix shows the joint model scores below an audio-only enhancer on waveform-alignment metrics (PESQ, STOI, SI-SDR), which implies that claims of audio superiority must be scoped to perceptual, no-reference measures; a reader should expect this trade-off in any generative audio restorer.
- The documented cross-window color drift suggests a concrete fix the paper leaves implicit: scheduled sampling that occasionally feeds the model its own restored frames as the next window's anchor would close the train/inference gap it identifies for first-frame anchoring.
- If the synthetic-to-real domain gap narrows with larger real paired data, the same fixed-prompt architecture transfers to other jointly degraded media, such as newsreels with voiceover or home movies with magnetic soundtracks, where per-clip captioning would be impractical.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents OmniVR, a joint audio-video generative restoration model for historical films, built by adapting a 22B-parameter multimodal diffusion transformer (LTX-2) from text-to-audio-video generation to audio-video-to-audio-video restoration. The contributions are a synthetic joint degradation pipeline, an architecture-preserving T2AV-to-AV2AV transition with prompt annealing, first-frame I2V anchoring with loss reweighting and waveform supervision, and a new benchmark, OmniVRBench, consisting of 200 real historical clips in a Real track and a Controlled track. Experiments report visual no-reference metrics, audio no-reference metrics, lip-sync metrics, full-reference metrics on the Controlled track, ablations, a human preference study, and external validation on the RTN old-film benchmark. The paper claims state-of-the-art results on all six visual metrics, best audio quality, and natural colorization.
Significance. If the central claims can be supported, the contribution is significant: OmniVR would be the first joint audio-video generative restoration model, and OmniVRBench would provide a useful evaluation protocol for a task that has previously been split between visual and audio communities. The paper has genuine strengths that should be credited: it reports external RTN validation for the visual component, full-reference LPIPS/DISTS/FVD on the Controlled track, fidelity probes such as face-identity consistency and OCR agreement, a human preference study, and unusually detailed appendices on training-data transparency, degradation parameter ranges, and known failure modes. The ablations in Table 4 are informative and mostly well designed. However, two load-bearing problems prevent the paper from being accepted in its current form. First, the headline claim that OmniVR 'achieves the best audio quality' is contradicted by the paper's own full-reference audio results (Appendix A.2, Table 7) and by its cascade baseline comparison (Appendix E.1, Table 11).
major comments (4)
- [Abstract; Section 4.4; Table 11; Table 7] The claim that OmniVR 'achieves the best audio quality' is contradicted by the paper's own numbers. In Table 11, the cascade baselines RealBasicVSR+VoiceFixer and MambaOFR+VoiceFixer beat OmniVR on DNSMOS on both tracks (2.96 and 2.79 versus 2.70 and 2.43) and beat OmniVR on FAD in three of four rows (6.24/8.29 versus 6.30/8.32 on Controlled/Real). In Table 7, OmniVR trails VoiceFixer on PESQ, STOI, and SI-SDR, and its STOI (0.510) is below that of the low-quality input (0.718). The paper itself concedes in Appendix A.2 that OmniVR is not the best on any waveform-aligned metric. The headline claim should be revised to state that OmniVR's audio strength lies in cross-modal consistency and lip synchronization rather than in general audio quality, or the main-text evaluation must be replaced with a protocol that actually supports the claim as stated.
- [Tables 2, 3, and 11] The audio evaluation is internally inconsistent. VoiceFixer alone is reported with DNSMOS 2.12 (Table 2, Controlled) and 2.21 (Table 3, Real), while the same VoiceFixer in cascade with RealBasicVSR or MambaOFR is reported with DNSMOS 2.96 and 2.79 on those same tracks, even though neither video restorer processes audio. These numbers cannot be simultaneously correct under a single evaluation protocol. The paper must specify the exact audio evaluation protocol (feature extraction, sampling, alignment, reference set), rerun all audio baselines under that protocol, and report them together in one table; otherwise the comparison is not internally consistent.
- [Section 3.2; Section 4.1; Appendix A; Section 5] The central evaluation is partly circular. The Controlled track is constructed by applying the same synthetic degradation operator D (Eq. 1, Section 3.2) that is used to generate training data, so the full-reference results in Table 6 measure the model's ability to invert its own training-time degradation distribution. The Real track has no references and is evaluated only with no-reference proxies, and the external RTN benchmark (Table 1) covers only visual quality and contains only three sequences. The paper acknowledges the domain gap in Section 5, but the main abstract and conclusion claims nevertheless rely on transfer to authentic archival footage. A concrete way to strengthen this would be to test on a held-out degradation distribution that differs from D, on real degraded/restored pairs if any can be obtained, or on a larger external benchmark that includes audio. Without such evidence, the claim that OmniVR generalizes to real historical films is not yet established.
- [Section 4.4; Appendix A.3] The main-text statement that OmniVR 'matches or exceeds the clean reference' on no-reference metrics (e.g., MUSIQ 71.17 vs. GT 67.49 in Table 2) is presented without the crucial caveat that the authors themselves provide in Appendix A.3: no-reference metrics reward generic clean-looking statistics and are not evidence of objective superiority over the ground truth. This is not a fatal flaw because the appendix is honest, but the main text should state the caveat at first use, since the phrase 'on par with the clean source' invites a stronger reading than the evidence supports.
minor comments (5)
- [Section 5; Appendix H] The first-frame anchor uses clean ground-truth tokens during training but self-generated restored frames during inference, a distribution shift that the authors themselves identify in Appendix H as the likely cause of cross-window color/texture drift. Since 'seamless long-video restoration' is a headline capability, this limitation should be stated explicitly in the main text rather than only in an appendix.
- [Table 5] The human preference study relies on 12 annotators; although bootstrap confidence intervals are reported, the main text should describe the annotation procedure (what the annotators saw, whether audio and video were presented together, and how ties were resolved).
- [Abstract; Section 4.4] The phrase 'all six visual metrics' is ambiguous; the abstract should specify that it refers to the no-reference visual metrics on OmniVRBench and RTN, not to full-reference metrics such as PSNR/SSIM/LPIPS.
- [Figure 3] The diagram contains duplicated labels 'LQ Gaussian Noise' and 'Sample' in both branches; the figure should be cleaned up to avoid confusion.
- [Section 3.3; Eq. (4)] The condition-noise range is written as 'rho_m ~ U(0.4,0.6)' in Eq. (4) but the caption and implementation details state rho=0.5 at inference; please confirm the training distribution is U(0.4,0.6) and not U(0.4,0.6) including endpoints, and state the inference value consistently.
Circularity Check
Controlled-track evaluation is partly self-referential because the benchmark's synthetic degradation operator D is the same operator used to create training pairs; independent RTN and real-track results prevent the circularity from being total.
-
fitted input called prediction
[Section 3.2 Eq. (1); Section 4.4 Controlled-track discussion; Appendix A opening paragraph]
"Because the Controlled track is built from clean sources via the synthetic degradation operator D (Eq. 1), a clean reference (Vh, Ah) exists for every clip there. ... This indicates that OmniVR does not merely invert synthetic degradation but produces perceptual quality on par with the clean source."
The training objective optimizes the model to invert pairs (Vl, Al) = D(Vh, Ah) from Eq. (1). The Controlled track is then constructed by applying that same operator D to clean sources, so the model's Controlled-track 'predictions' are evaluated on the exact distribution it was trained to invert. The claim that this result shows OmniVR 'does not merely invert synthetic degradation' is therefore not supported by the Controlled track: an in-distribution test cannot distinguish inversion of D from genuine generalization. The circularity is partial because the Real Historical track and the external RTN benchmark are not generated by D and provide independent evidence.
full rationale
The paper's central derivation chain is not mathematically circular: the conditional-generation formulation, the LoRA adaptation of LTX-2, the prompt-annealing curriculum, and the first-frame anchoring are all independently motivated and evaluated with ablations. The main circularity risk is evaluation-level: the Controlled track of OmniVRBench is produced by the same synthetic degradation operator D used to generate training data, so Controlled-track superiority is partly a measure of in-distribution fit rather than a standalone demonstration of generalization. This is partially mitigated by the Real Historical track (no synthetic D), the external RTN visual benchmark, and the human preference study, as the paper itself notes the proxy-based nature of real-track evaluation in Section 5. Separately, the 'best audio quality' claim is weakened by the paper's own full-reference audio table (OmniVR is not best on PESQ/STOI/SI-SDR) and by Appendix Table 11, where cascade baselines exceed OmniVR on DNSMOS; however, that is a metric-selection and internal-consistency issue, not a circular-derivation step, so it does not further raise the circularity score.
Assumptions & free parameters
free parameters (6)
- loss reweighting coefficients (alpha_v, alpha_a, beta_v, beta_a) =
1.0, 1.0, 0.5, 0.5
- STFT loss weight (lambda_stft) =
3e-3
- CFG scale =
3.0
- condition noise level (rho) =
U(0.4,0.6) training, 0.5 inference
- first-frame anchor probability (p_ff) =
0.5
- prompt annealing schedule =
caption prob decays from 1.0 to 0 over first 30% of training
assumptions (5)
- domain assumption The pretrained LTX-2 22B T2AV backbone is a suitable base for AV2AV restoration and retains its generative prior under LoRA adaptation.
- domain assumption The synthetic joint degradation pipeline D faithfully approximates real historical film degradations.
- domain assumption No-reference metrics (MUSIQ, NIQE, DNSMOS, FAD) are valid proxies for restoration quality on historical films.
- domain assumption Joint denoising with cross-modal gates improves both modalities compared to independent restoration.
- domain assumption Flow matching training with frozen VAEs and LoRA converges to a functional restorer under the composite loss.
Cite this review
Pith. "Pith review of OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films." pith.science (2026). https://pith.science/paper/GJGTIHO5
@misc{pith2026260804224,
author = {Pith},
title = {Pith review of: OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films},
year = {2026},
howpublished = {\url{https://pith.science/paper/GJGTIHO5}},
note = {Machine review of arXiv:2608.04224}
}
read the original abstract
Historical films suffer from co-occurring visual and audio degradations---blur, noise, flicker, hiss, clipping, and dropout---yet existing methods restore each modality independently, leaving quality gaps and cross-modal inconsistency. We present OmniVR, the first joint audio-video generative restoration model. Built upon a 22B-parameter audio-video generation backbone, OmniVR formulates restoration as conditional generation within a unified multimodal DiT: the low-quality video and audio are encoded as latent conditions, combined with a fixed restoration prompt, and jointly denoised to recover visual structure, temporal motion, and acoustic detail under one coordinated objective. Three key designs enable this adaptation: (1) a joint audio-video degradation pipeline that simulates real old-film characteristics from Internet-collected data; (2) an architecture-preserving text-to-audio-video (T2AV) to audio-video-to-audio-video (AV2AV) transition with prompt annealing that maximally retains the generative prior; and (3) first-frame image-to-video (I2V) anchoring with loss reweighting and waveform supervision for long-video extrapolation and audio fidelity. We also propose OmniVRBench, the first benchmark that evaluates audio-video restoration across visual quality, audio quality, temporal consistency, and audio-visual synchrony on 200 real historical clips. OmniVR surpasses all prior methods on all six visual metrics, achieves the best audio quality, and produces natural colorization---the first method to jointly address all three aspects. Code and weights will be publicly released. Project Page: https://xin1u.github.io/OminiVR_PAGE/
Figures
Figures from the paper (17 more)
Reference graph
Works this paper leans on
-
[1]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
Making Old Film Great Again: Degradation-Aware State Space Model for Old Film Restoration , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
-
[2]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
-
[3]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
Investigating Tradeoffs in Real-World Video Super-Resolution , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
-
[4]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
VisualVoice: Audio-Visual Speech Separation with Cross-Modal Consistency , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
-
[5]
arXiv preprint arXiv:2601.03233 , year=
LTX-2: Efficient Joint Audio-Visual Foundation Model , author=. arXiv preprint arXiv:2601.03233 , year=
-
[6]
Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , volume=
Denoising diffusion probabilistic models , author=. Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , volume=
-
[7]
arXiv preprint arXiv:2207.12598 , year=
Classifier-Free Diffusion Guidance , author=. arXiv preprint arXiv:2207.12598 , year=
-
[8]
Proceedings of the International Conference on Learning Representations (ICLR) , year=
LoRA: Low-Rank Adaptation of Large Language Models , author=. Proceedings of the International Conference on Learning Representations (ICLR) , year=
Show all 54 references
-
[9]
arXiv preprint arXiv:2411.13503 , year=
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models , author=. arXiv preprint arXiv:2411.13503 , year=
-
[10]
ACM Transactions on Graphics (TOG) , volume=
DeepRemaster: Temporal Source-Reference Attention Networks for Comprehensive Video Enhancement , author=. ACM Transactions on Graphics (TOG) , volume=. 2019 , publisher=
2019
-
[11]
Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , year=
MUSIQ: Multi-scale Image Quality Transformer , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , year=
-
[12]
Kilgour, Kevin and Zuluaga, Mauricio and Roblek, Dominik and Sharifi, Matthew , booktitle=. Fr
-
[13]
Proceedings of the International Conference on Machine Learning (ICML) , year=
VideoPoet: A Large Language Model for Zero-Shot Video Generation , author=. Proceedings of the International Conference on Machine Learning (ICML) , year=
-
[14]
Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , year=
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis , author=. Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , year=
-
[15]
Proceedings of the International Conference on Learning Representations (ICLR) , year=
DiffWave: A Versatile Diffusion Model for Audio Synthesis , author=. Proceedings of the International Conference on Learning Representations (ICLR) , year=
-
[16]
arXiv preprint arXiv:1811.02508 , year=
SDR--Half-Baked or Well Done? , author=. arXiv preprint arXiv:1811.02508 , year=
-
[17]
Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , year=
Recurrent Video Restoration Transformer with Guided Deformable Attention , author=. Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , year=
-
[18]
Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , year=
VRT: A Video Restoration Transformer , author=. Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , year=
-
[19]
Proceedings of the European Conference on Computer Vision (ECCV) , year=
DiffBIR: Toward Blind Image Restoration with Generative Diffusion Prior , author=. Proceedings of the European Conference on Computer Vision (ECCV) , year=
-
[20]
arXiv preprint arXiv:2109.13731 , year=
VoiceFixer: Toward General Speech Restoration with Neural Vocoder , author=. arXiv preprint arXiv:2109.13731 , year=
-
[21]
Proceedings of the International Conference on Learning Representations (ICLR) , year=
Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow , author=. Proceedings of the International Conference on Learning Representations (ICLR) , year=
-
[22]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
-
[23]
Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , year=
Scalable Diffusion Models with Transformers , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , year=
-
[24]
arXiv preprint arXiv:2410.13720 , year=
Movie Gen: A Cast of Media Foundation Models , author=. arXiv preprint arXiv:2410.13720 , year=
-
[25]
Proceedings of the ACM International Conference on Multimedia (ACM MM) , year=
A Lip Sync Expert Is All You Need for Speech to Lip Generation in the Wild , author=. Proceedings of the ACM International Conference on Multimedia (ACM MM) , year=
-
[26]
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year=
DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors , author=. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year=
-
[27]
Journal of the Audio Engineering Society , volume=
Perceptual Evaluation of Speech Quality (PESQ): The New ITU Standard for End-to-End Speech Quality Assessment Part I--Time-Delay Compensation , author=. Journal of the Audio Engineering Society , volume=
-
[28]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=
High-resolution image synthesis with latent diffusion models , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=
-
[29]
IEEE Transactions on Pattern Analysis and Machine Intelligence , volume=
Image super-resolution via iterative refinement , author=. IEEE Transactions on Pattern Analysis and Machine Intelligence , volume=
-
[30]
Proceedings of the International Conference on Learning Representations (ICLR) , year=
Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction , author=. Proceedings of the International Conference on Learning Representations (ICLR) , year=
-
[31]
IEEE Transactions on Audio, Speech, and Language Processing , volume=
An Algorithm for Intelligibility Prediction of Time-Frequency Weighted Noisy Speech , author=. IEEE Transactions on Audio, Speech, and Language Processing , volume=
-
[32]
arXiv preprint arXiv:1812.01717 , year=
Towards Accurate Generative Models of Video: A New Metric and Challenges , author=. arXiv preprint arXiv:1812.01717 , year=
-
[33]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
Bringing Old Photos Back to Life , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
-
[34]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
Bringing Old Films Back to Life , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
-
[35]
IEEE transactions on image processing , volume=
Image quality assessment: from error visibility to structural similarity , author=. IEEE transactions on image processing , volume=. 2004 , publisher=
2004
-
[36]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , year=
EDVR: Video Restoration with Enhanced Deformable Convolutional Networks , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , year=
-
[37]
Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) , year=
Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) , year=
-
[38]
Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , year=
Exploring CLIP for Assessing the Look and Feel of Images , author=. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , year=
-
[39]
International Journal of Computer Vision (IJCV) , year=
Exploiting Diffusion Prior for Real-World Image Super-Resolution , author=. International Journal of Computer Vision (IJCV) , year=
-
[40]
Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , year=
Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , year=
-
[41]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
VideoGigaGAN: Towards Detail-Rich Video Super-Resolution , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
-
[42]
arXiv preprint arXiv:2605.24652 , year=
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models , author=. arXiv preprint arXiv:2605.24652 , year=
-
[43]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
-
[44]
Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages=
The unreasonable effectiveness of deep features as a perceptual metric , author=. Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages=
-
[45]
arXiv preprint arXiv:2607.00726 , year=
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization , author=. arXiv preprint arXiv:2607.00726 , year=
-
[46]
European Conference on Computer Vision (ECCV) , year=
ColorMNet: A Memory-Based Deep Spatial-Temporal Feature Propagation Network for Video Colorization , author=. European Conference on Computer Vision (ECCV) , year=
-
[47]
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , year=
BiSTNet: Semantic Image Prior Guided Bidirectional Temporal Feature Fusion for Deep Exemplar-Based Video Colorization , author=. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , year=
-
[48]
IEEE/CVF International Conference on Computer Vision (ICCV) , year=
DDColor: Towards Photo-Realistic Image Colorization via Dual Decoders , author=. IEEE/CVF International Conference on Computer Vision (ICCV) , year=
-
[49]
Human Vision and Electronic Imaging VIII (SPIE) , volume=
Measuring Colorfulness in Natural Images , author=. Human Vision and Electronic Imaging VIII (SPIE) , volume=
-
[50]
IEEE Signal Processing Letters , volume=
Making a ``Completely Blind'' Image Quality Analyzer , author=. IEEE Signal Processing Letters , volume=
-
[51]
IEEE Transactions on Image Processing , volume=
No-Reference Image Quality Assessment in the Spatial Domain , author=. IEEE Transactions on Image Processing , volume=
-
[52]
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , year=
MANIQA: Multi-Dimension Attention Network for No-Reference Image Quality Assessment , author=. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , year=
-
[53]
IEEE Transactions on Image Processing , year=
TOPIQ: A Top-Down Approach from Semantics to Distortions for Image Quality Assessment , author=. IEEE Transactions on Image Processing , year=
-
[54]
Liu, Haohe and Kong, Qiuqiang and Tian, Qiao and Zhao, Yan and Wang, DeLiang and Huang, Chuanzeng and Wang, Yuxuan , booktitle=
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.