REVIEW 2 major objections 5 minor 65 references
Augmented Deep Contexts for Spatially Embedded Video Coding
T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read SEVC claims that embedding a 4×-downsampled spatial reference lets neural video codecs beat temporal-only prediction, cutting bitrate by 11.9% at equal quality.
desk verdict A credible spatially scalable neural video codec with honest 11.9% bitrate savings; the main open question is the undisclosed fine-tuning subset selection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Motion and Feature Co-Augmentation (MFCA) module: a multi-scale, multi-stage loop in which each Augment Stage first sharpens the low-resolution motion vectors by adding a residual predicted from temporal features and the current spatial feature, then warps the temporal feature with the sharpened motion and uses it to refine the spatial feature. This progressive refinement generates the hybrid spatial-temporal contexts \($C_t^{1}$, $C_t^{2}$, $C_t^{3}$\). The second load-bearing device is the spatial-guided latent prior: the spatial latent \(\hat{y}_t^b\), upsampled to full resolution, serves as the query that aligns \(\hat{y}_{t-1}, \hat{y}_{t-2}, \hat{y}_{t-3}\) through Transformer blocks, replacing the single misaligned prior \(\hat{y}_{t-1}\). A joint spatial-temporal optimization with a small reconstruction constraint on the base layer lets the network learn how many bits the low-resolution stream should spend, rather than fixing that by hand.
What would settle it
Encode a fast-motion or object-appearance sequence with SEVC and with a version of SEVC in which the low-resolution spatial branch is replaced by a constant input (the average low-resolution color) while the base bits are reallocated to the full-resolution layer; if BD-rate over DCVC-DC does not degrade, the spatial reference is not the source of the gain. A second check is to blur the low-resolution input before encoding and see whether the reported 11.9% saving shrinks.
Extended reading notes
Core claim
The central discovery is that a lossy low-resolution reconstruction of the current frame can be converted into a high-value prediction asset rather than a separate bitstream burden. SEVC's MFCA module co-augments base motion vectors and a spatial feature by alternating residual prediction with temporal-feature alignment across three scales, yielding hybrid contexts that describe regions where motion estimation or temporal reference is unreliable. In parallel, the low-resolution latent, upsampled and used as a transformer query, aligns and fuses the previous three latent representations into a spatial-guided prior for the entropy model. The paper reports average BD-rate savings of 11.9% over DCVC-FM at an intra period of -1 and 8.3% over DCVC-DC at an intra period of 32, with the largest margins on sequences containing large motions or emerging objects.
Load-bearing premise
The design banks on the idea that a 4×-smaller, compressed copy of the current frame retains enough real spatial detail to guide full-resolution prediction; if that copy is too blurry or too lossy, the motion-and-feature augmentation and the latent prior query have nothing useful to add.
Editorial extensions
If this is right
- On sequences with large motion or newly appearing objects, SEVC should show its largest BD-rate advantage over temporal-only codecs, since that is the regime the spatial branch is designed to repair.
- A SEVC bitstream can be partially decoded into a low-resolution video, so fast preview and skimming become possible from the same compressed representation.
- The spatial-guided prior should make rate estimation more accurate for frames that differ strongly from the previous frame, lowering bitrate at equal quality.
- Base-layer bit allocation can be learned end-to-end, removing the need to hand-tune quality ratios between the low- and full-resolution layers.
- The spatial-embedding strategy is an augmentation over a base codec, so improved future temporal-only codecs could inherit the same gain by being wrapped in the same structure.
Reading between the lines
- A stress test on scene cuts and content swaps would likely amplify the reported gains, because temporal references become nearly useless there and the spatial branch carries almost the entire prediction load; the paper's tested sets contain such moments but do not isolate them.
- The 4× downsampling factor is chosen, not proven optimal; a smaller factor would send more spatial detail at higher base cost, so the best trade-off may shift with resolution, bitrate, and content, and locating the optimum would require a sweep the paper does not report.
- Because the base layer is a low-resolution decodable stream, SEVC has a natural fit for adaptive streaming use cases where a client first requests the low-resolution layer and later upgrades; the paper does not explore that deployment path.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SEVC embeds a low-resolution base codec (DCVC-DC style) into a full-resolution learned video codec: the input frame is 4x downsampled and compressed, yielding base motion vectors, a spatial feature, and a spatial latent representation. These spatial references are then augmented with temporal references in two ways: the Motion and Feature Co-Augmentation (MFCA) module progressively refines base MVs and the spatial feature to produce hybrid spatial-temporal contexts, and a spatial-guided Transformer-based latent prior aligns and fuses multiple past latents. A joint spatial-temporal optimization loss adjusts base-layer bit allocation. Experiments report BD-Rate savings over VTM-13.2 and several learned codecs, including an 11.9% additional saving over DCVC-FM at IP=-1, together with targeted evidence on large-motion and emerging-object sequences. Code and trained models are released.
Significance. If the empirical claims hold, SEVC is a meaningful contribution: it transfers the spatial-reference idea from scalable and super-resolution coding into conditional NVC feature-space context generation, and it includes the base-layer bits in the total rate, so the reported savings are not obtained by hiding the extra low-resolution bitstream. The ablation structure is sensible and supports the individual design choices, and the public code and models strengthen reproducibility. The main caveat is that the headline result is produced by a fine-tuning stage whose training subset is not specified, which makes the central number hard to verify; the large-motion/emerging-object claim also rests on a very small set of sequences. These issues are fixable with disclosure or additional experiments, so the contribution is potentially publishable in a major revision.
major comments (2)
- [4.1 (Training Setup)] The joint optimization is conducted on "a selected subset of 9000 sequences from the original videos of the Vimeo-90k dataset," but the selection criterion is never given. Because the headline 11.9% BD-Rate improvement and the large-motion/emerging-object results are outcomes of this fine-tuning stage, the subset choice is load-bearing: if it was chosen to emphasize large-motion or emerging-object content, or to resemble the test sequences in Table 3, the reported gains could be inflated. Please state the exact selection criterion, specify whether the selection was random, and release the list of sequence indices; if the selection was content-based, repeat the fine-tuning on a random subset and report both sets of results.
- [4.2 and Table 3] The claim that SEVC "effectively alleviates the limitations in handling large motions or emerging objects" is supported by only three named sequences (USTC BicycleDriving, videoSRC21, BasketballDrive) with no a priori protocol for choosing them and no error bars. One of the three comes from USTC-TD, a dataset introduced by the same group. To make the claim convincing, the paper should either provide a systematic evaluation over a larger set of sequences with a defined threshold or annotation for "large motion" and "emerging object," or substantially soften the claim to a qualitative observation. Without this, the strong wording in the abstract and conclusion exceeds what the evidence supports.
minor comments (5)
- [Various] There are several typos: Section 2.1 has "the the superior potential," Section 3.1 has "persepecitive," and the supplementary material has "Euqation" and "Architechture." These should be corrected.
- [3.3] The statement that the hyperprior is discarded "due to similar characteristics of hyper encoder/decoder and our base codec" is asserted without an ablation. Please either add an experiment keeping the hyperprior or soften the claim.
- [4.3 (Table 5)] The bullet notation in Table 5 is ambiguous: it is not immediately clear which components are present in M1 and M4, especially because the text says M4 discards the spatial latent. Please spell out each baseline in words or use explicit check marks with a legend.
- [Equations (4) and (5)] Define R_t explicitly as the total bitrate including the base-layer bits. The text implies this, but a formal definition would prevent readers from misinterpreting the BD-Rate comparisons as excluding the low-resolution bitstream.
- [Tables 1-6] No error bars, confidence intervals, or repeated-seed results are reported. Some differences are small (for example, 1.5% between M1 and M2 in Table 5), so a statement about variance or a release of per-sequence numbers would strengthen the empirical claims.
Circularity Check
No circular derivation found; empirical benchmark claims stand on held-out tests, with a training-subset disclosure caveat.
full rationale
I walked the derivation chain. Equation (1), H(X) = H(X^b) + H(X|X^b), is a standard information-theoretic identity used only as motivation, and the supplementary derivation is correct; it does not define SEVC's performance in terms of itself. The reported BD-Rate gains (Tables 1-3) are measured on separate test sets (HEVC B-E, MCL-JCV, UVG, USTC-TD) against VTM, with the base-layer bits included in the total rate, so the 11.9% claim is not a fitted input renamed as a prediction. Hyperparameters such as lambda and w_l are fixed regularization/loss weights selected by ablations, not inverse-fitted to test RD values. Self-citations (e.g., [6] LSSVC, [32] USTC-TD, [51] Sheng-2024) are design or test-set references and are not used as a uniqueness theorem or to forbid alternatives; the base codec DCVC-DC [28] is external prior work. The one flagged issue is not circular: Section 4.1 says joint optimization is 'conducted on a selected subset of 9000 sequences from the original videos of the Vimeo-90k dataset' without specifying the selection criterion. This is a training-data selection and reproducibility concern that could bias the claimed gains, but it does not make any equation reduce to its own input or any fitted parameter masquerade as a prediction. Hence the low circularity score.
Assumptions & free parameters
free parameters (5)
- lambda (RD tradeoff) =
{50, 95, 200, 400}
- wl (base reconstruction weight) =
0.05 in ablations; final value not stated
- number of augment stages =
2
- number of temporal latent representations =
3
- downsampling factor =
4x
assumptions (5)
- standard math H(X)=H(Xb)+H(X|Xb) for deterministic downsampling
- domain assumption DCVC-DC reconstruction provides reliable spatial references
- ad hoc to paper Transformer with spatial latent as query aligns temporal latents
- domain assumption Removing the hyperprior does not degrade RD performance
- domain assumption Bicubic downsampling models the spatial reference degradation
Cite this review
Pith. "Pith review of Augmented Deep Contexts for Spatially Embedded Video Coding." pith.science (2026). https://pith.science/paper/IXLOX62Y
@misc{pith2026250505309,
author = {Pith},
title = {Pith review of: Augmented Deep Contexts for Spatially Embedded Video Coding},
year = {2026},
howpublished = {\url{https://pith.science/paper/IXLOX62Y}},
note = {Machine review of arXiv:2505.05309}
}
read the original abstract
Most Neural Video Codecs (NVCs) only employ temporal references to generate temporal-only contexts and latent prior. These temporal-only NVCs fail to handle large motions or emerging objects due to limited contexts and misaligned latent prior. To relieve the limitations, we propose a Spatially Embedded Video Codec (SEVC), in which the low-resolution video is compressed for spatial references. Firstly, our SEVC leverages both spatial and temporal references to generate augmented motion vectors and hybrid spatial-temporal contexts. Secondly, to address the misalignment issue in latent prior and enrich the prior information, we introduce a spatial-guided latent prior augmented by multiple temporal latent representations. At last, we design a joint spatial-temporal optimization to learn quality-adaptive bit allocation for spatial references, further boosting rate-distortion performance. Experimental results show that our SEVC effectively alleviates the limitations in handling large motions or emerging objects, and also reduces 11.9% more bitrate than the previous state-of-the-art NVC while providing an additional low-resolution bitstream. Our code and model are available at https://github.com/EsakaK/SEVC.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
https://vcgit.hhi.fraunhofer.de/ jvet/VVCSoftware_VTM
VTM-17.0. https://vcgit.hhi.fraunhofer.de/ jvet/VVCSoftware_VTM . Accessed July 28, 2024. 6, 12
work page 2024
-
[2]
https://github.com/sanghyun- son/bicubic_pytorch
bicubic-pytorch. https://github.com/sanghyun- son/bicubic_pytorch. Accessed July 28, 2024. 6
work page 2024
-
[3]
Scale-space flow for end-to-end optimized video compression
Eirikur Agustsson, David Minnen, Nick Johnston, Johannes Balle, Sung Jin Hwang, and George Toderici. Scale-space flow for end-to-end optimized video compression. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8503–8512, 2020. 2
2020
-
[4]
Hierarchical B-frame Video Coding Using Two- Layer CANF without Motion Coding
David Alexandre, Hsueh-Ming Hang, and Wen-Hsiao Peng. Hierarchical B-frame Video Coding Using Two- Layer CANF without Motion Coding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10249–10258, 2023. 3
work page 2023
-
[5]
Variational image compres- sion with a scale hyperprior
Johannes Ball ´e, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. Variational image compres- sion with a scale hyperprior. In International Conference on Learning Representations (ICLR) , pages 4961–5007, 2018. 4
work page 2018
-
[6]
LSSVC: A Learned Spatially Scalable Video Coding Scheme
Yifan Bian, Xihua Sheng, Li Li, and Dong Liu. LSSVC: A Learned Spatially Scalable Video Coding Scheme. IEEE Transactions on Image Processing: a publication of the IEEE Signal Processing Society, 33:3314–3327, 2024. 2
work page 2024
-
[7]
Common test conditions and software reference configurations
Frank Bossen et al. Common test conditions and software reference configurations. JCTVC-L1100, 12(7):1, 2013. 6, 7
work page 2013
-
[8]
Overview of SHVC: Scalable Extensions of the High Efficiency Video Coding Standard
Jill M Boyce, Yan Ye, Jianle Chen, and Adarsh K Ramasub- ramonian. Overview of SHVC: Scalable Extensions of the High Efficiency Video Coding Standard. IEEE Transactions on Circuits and Systems for Video Technology, 26(1):20–34,
Show all 65 references
-
[9]
Overview of the Versatile Video Coding (VVC) Standard and its Applica- tions
Benjamin Bross, Ye-Kui Wang, Yan Ye, Shan Liu, Jianle Chen, Gary J Sullivan, and Jens-Rainer Ohm. Overview of the Versatile Video Coding (VVC) Standard and its Applica- tions. IEEE Transactions on Circuits and Systems for Video Technology, 31(10):3736–3764, 2021. 1
2021
-
[10]
BasicVSR: The Search for Essential Components in Video Super-Resolution and Beyond
Kelvin CK Chan, Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. BasicVSR: The Search for Essential Components in Video Super-Resolution and Beyond. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 4947–4956, 2021. 1, 3
2021
-
[11]
BasicVSR++: Improving Video Super- Resolution with Enhanced Propagation and Alignment
Kelvin CK Chan, Shangchen Zhou, Xiangyu Xu, and Chen Change Loy. BasicVSR++: Improving Video Super- Resolution with Enhanced Propagation and Alignment. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR) , pages 5972–5981,
-
[12]
NeRV: Neural Representations for Videos
Hao Chen, Bo He, Hanyu Wang, Yixuan Ren, Ser Nam Lim, and Abhinav Shrivastava. NeRV: Neural Representations for Videos. Advances in Neural Information Processing Systems (NeurIPS), 34:21557–21568, 2021. 2
2021
-
[13]
HNeRV: A Hybrid Neural Representation for Videos
Hao Chen, Matthew Gwilliam, Ser-Nam Lim, and Abhinav Shrivastava. HNeRV: A Hybrid Neural Representation for Videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 10270–10279, 2023. 2
2023
-
[14]
Elements of information theory
Thomas M Cover. Elements of information theory . John Wiley & Sons, 1999. 3, 13
1999
-
[15]
Neural Inter-Frame Com- pression for Video Coding
Abdelaziz Djelouah, Joaquim Campos, Simone Schaub- Meyer, and Christopher Schroers. Neural Inter-Frame Com- pression for Video Coding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 6421–6429, 2019. 2
2019
-
[16]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, G Heigold, S Gelly, et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on...
2020
-
[17]
Asymmetric numeral systems: entropy coding combining speed of Huffman coding with compression rate of arithmetic coding
Jarek Duda. Asymmetric numeral systems: entropy coding combining speed of Huffman coding with compression rate of arithmetic coding. arXiv preprint arXiv:1311.2540, 2013. 4
2013 arXiv
-
[18]
Comparison of the H.263 and H.261 video compression stan- dards
Bernd Girod, Eckehard G Steinbach, and Niko Faerber. Comparison of the H.263 and H.261 video compression stan- dards. In Standards and Common Interfaces for Video Infor- mation Systems: A Critical Review , pages 230–248. SPIE,
-
[19]
CANF-VC: Conditional Augmented Normalizing Flows for Video Compression
Yung-Han Ho, Chih-Peng Chang, Peng-Yu Chen, Alessan- dro Gnutti, and Wen-Hsiao Peng. CANF-VC: Conditional Augmented Normalizing Flows for Video Compression. In European Conference on Computer Vision (ECCV) , pages 207–223. Springer, 2022. 2, 5
2022
-
[20]
FVC: A New Framework towards Deep Video Compression in Feature Space
Zhihao Hu, Guo Lu, and Dong Xu. FVC: A New Framework towards Deep Video Compression in Feature Space. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 1502–1511, 2021. 1, 2, 3
2021
-
[21]
Coarse-to-fine Deep Video Coding with Hyperprior-guided Mode Prediction
Zhihao Hu, Guo Lu, Jinyang Guo, Shan Liu, Wei Jiang, and Dong Xu. Coarse-to-fine Deep Video Coding with Hyperprior-guided Mode Prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5921–5930, 2022. 2, 3, 5
2022
-
[22]
Sparse Point Clouds Assisted Learned Image Compression
Yiheng Jiang, Haotian Zhang, Li Li, Dong Liu, and Zhu Li. Sparse Point Clouds Assisted Learned Image Compression. arXiv preprint arXiv:2412.15752, 2024. 2
2024 arXiv
-
[23]
Video Super-Resolution With Con- volutional Neural Networks
Kappeler, Armin and Yoo, Seunghwan and Dai, Qiqin and Katsaggelos, Aggelos K. Video Super-Resolution With Con- volutional Neural Networks. IEEE Transactions on Compu- tational Imaging, 2(2):109–122, 2016. 1, 3
2016
-
[24]
Optical Flow and Mode Se- lection for Learning-based Video Coding
Th ´eo Ladune, Pierrick Philippe, Wassim Hamidouche, Lu Zhang, and Olivier D ´eforges. Optical Flow and Mode Se- lection for Learning-based Video Coding. In 2020 IEEE 22nd International Workshop on Multimedia Signal Process- ing (MMSP), pages 1–6. IEEE, 2020. 2
2020
-
[25]
Conditional Coding and Vari- able Bitrate for Practical Learned Video Coding
Th ´eo Ladune, Pierrick Philippe, Wassim Hamidouche, Lu Zhang, and Olivier D´eforges. Conditional Coding and Vari- able Bitrate for Practical Learned Video Coding. arXiv preprint arXiv:2104.09103, 2021
2021 arXiv
-
[26]
Deep Contextual Video Com- pression
Jiahao Li, Bin Li, and Yan Lu. Deep Contextual Video Com- pression. Advances in Neural Information Processing Sys- tems (NeurIPS), 34:18114–18125, 2021. 2, 3, 4, 5, 6 9
2021
-
[27]
Hybrid Spatial-Temporal En- tropy Modelling for Neural Video Compression
Jiahao Li, Bin Li, and Yan Lu. Hybrid Spatial-Temporal En- tropy Modelling for Neural Video Compression. InProceed- ings of the 30th ACM International Conference on Multime- dia (ACM MM), pages 1503–1511, 2022. 1, 2, 3, 4, 6, 7, 8, 12, 14
2022
-
[28]
Neural Video Compression with Diverse Contexts
Jiahao Li, Bin Li, and Yan Lu. Neural Video Compression with Diverse Contexts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22616–22626, 2023. 1, 2, 3, 4, 5, 6, 7, 12, 13, 14
2023
-
[29]
Neural Video Compression with Feature Modulation
Jiahao Li, Bin Li, and Yan Lu. Neural Video Compression with Feature Modulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26099–26108, 2024. 1, 2, 3, 4, 6, 7, 13, 14
2024
-
[30]
E-NeRV: Expedite Neural Video Representation with Disentangled Spatial-Temporal Con- text
Zizhang Li, Mengmeng Wang, Huaijin Pi, Kechun Xu, Jian- biao Mei, and Yong Liu. E-NeRV: Expedite Neural Video Representation with Disentangled Spatial-Temporal Con- text. In European Conference on Computer Vision (ECCV), pages 267–284. Springer, 2022. 2
2022
-
[31]
Uniformly Accelerated Motion Model for Inter Prediction
Zhuoyuan Li, Yao Li, Chuanbo Tang, Li Li, Dong Liu, and Feng Wu. Uniformly Accelerated Motion Model for Inter Prediction. arXiv preprint arXiv:2407.11541, 2024. 1
2024 arXiv
-
[32]
USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020s
Zhuoyuan Li, Junqi Liao, Chuanbo Tang, Haotian Zhang, Yuqi Li, Yifan Bian, Xihua Sheng, Xinmin Feng, Yao Li, Changsheng Gao, et al. USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020s. arXiv preprint arXiv:2409.08481, 2024. 6, 7, 15
2024 arXiv
-
[33]
Object Segmentation-Assisted Inter Predic- tion for Versatile Video Coding
Zhuoyuan Li, Zikun Yuan, Li Li, Dong Liu, Xiaohu Tang, and Feng Wu. Object Segmentation-Assisted Inter Predic- tion for Versatile Video Coding. IEEE Transactions on Broadcasting, 70(4):1236–1253, 2024. 1
2024
-
[34]
SwinIR: Image Restoration Using Swin Transformer
Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. SwinIR: Image Restoration Using Swin Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1833–1844, 2021. 5, 12, 13
2021
-
[35]
M- LVC: Multiple Frames Prediction for Learned Video Com- pression
Jianping Lin, Dong Liu, Houqiang Li, and Feng Wu. M- LVC: Multiple Frames Prediction for Learned Video Com- pression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 3546–3554, 2020. 3
2020
-
[36]
Neural Video Coding using Mul- tiscale Motion Compensation and Spatiotemporal Context Model
Haojie Liu, Ming Lu, Zhan Ma, Fan Wang, Zhihuang Xie, Xun Cao, and Yao Wang. Neural Video Coding using Mul- tiscale Motion Compensation and Spatiotemporal Context Model. IEEE Transactions on Circuits and Systems for Video Technology, 31(8):3182–3196, 2020. 2
2020
-
[37]
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 10012–10022, 2021. 2, 5
2021
-
[38]
DVC: An End-to-end Deep Video Compression Framework
Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, and Zhiyong Gao. DVC: An End-to-end Deep Video Compression Framework. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11006–11015, 2019. 2, 3, 5
2019
-
[39]
An End-to-End Learning Framework for Video Compression
Guo Lu, Xiaoyun Zhang, Wanli Ouyang, Li Chen, Zhiyong Gao, and Dong Xu. An End-to-End Learning Framework for Video Compression. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(10):3292–3308, 2020. 2, 3, 5
2020
-
[40]
Don’t Blame the Elbo! A Linear Vae Perspec- tive on Posterior Collapse
James Lucas, George Tucker, Roger B Grosse, and Moham- mad Norouzi. Don’t Blame the Elbo! A Linear Vae Perspec- tive on Posterior Collapse. Advances in Neural Information Processing Systems (NeurIPS), 32, 2019. 5
2019
-
[41]
Uncertainty-Aware Deep Video Compression With Ensembles
Wufei Ma, Jiahao Li, Bin Li, and Yan Lu. Uncertainty-Aware Deep Video Compression With Ensembles. IEEE Transac- tions on Multimedia, 26:7863–7872, 2024. 1, 3
2024
-
[42]
Spatial scalability with VVC: coding performance and complexity
Gwenaelle Marquant, Charles Salmon-Legagneur, Fabrice Urban, and Philippe de Lagrange. Spatial scalability with VVC: coding performance and complexity. In Applications of Digital Image Processing XLV, pages 10–15. SPIE, 2022. 5
2022
-
[43]
VCT: A Video Compression Transformer
Fabian Mentzer, George Toderici, David Minnen, Sung-Jin Hwang, Sergi Caelles, Mario Lucic, and Eirikur Agustsson. VCT: A Video Compression Transformer. arXiv preprint arXiv:2206.07307, 2022. 3
2022 arXiv
-
[44]
UVG Dataset: 50/120fps 4K Sequences for Video Codec Analysis and Development
Alexandre Mercat, Marko Viitanen, and Jarno Vanne. UVG Dataset: 50/120fps 4K Sequences for Video Codec Analysis and Development. In Proceedings of the 11th ACM Multi- media Systems Conference (ACM MMSys) , pages 297–302,
-
[45]
Optical Flow Estima- tion using A Spatial Pyramid Network
Anurag Ranjan and Michael J Black. Optical Flow Estima- tion using A Spatial Pyramid Network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4161–4170, 2017. 7
2017
-
[46]
H.263: Video coding for low-bit-rate commu- nication
Karel Rijkse. H.263: Video coding for low-bit-rate commu- nication. IEEE Communications magazine , 34(12):42–45,
-
[47]
Learned Video Compression
Oren Rippel, Sanjay Nair, Carissa Lew, Steve Branson, Alexander G Anderson, and Lubomir Bourdev. Learned Video Compression. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV) , pages 3454–3463, 2019. 2
2019
-
[48]
ELF-VC: Effi- cient Learned Flexible-Rate Video Coding
Oren Rippel, Alexander G Anderson, Kedar Tatwawadi, San- jay Nair, Craig Lytle, and Lubomir Bourdev. ELF-VC: Effi- cient Learned Flexible-Rate Video Coding. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 14479–14488, 2021. 2
2021
-
[49]
Overview of the Scalable Video Coding Extension of the H.264/A VC Standard
Heiko Schwarz, Detlev Marpe, and Thomas Wiegand. Overview of the Scalable Video Coding Extension of the H.264/A VC Standard. IEEE Transactions on Circuits and Systems for Video Technology, 17(9):1103–1120, 2007. 5
2007
-
[50]
Temporal Context Mining for Learned Video Compression
Xihua Sheng, Jiahao Li, Bin Li, Li Li, Dong Liu, and Yan Lu. Temporal Context Mining for Learned Video Compression. IEEE Transactions on Multimedia, 25:7311–7322, 2022. 1, 2, 3, 4, 5, 6, 12
2022
-
[51]
Spatial De- composition and Temporal Fusion Based Inter Prediction for Learned Video Compression
Xihua Sheng, Li Li, Dong Liu, and Houqiang Li. Spatial De- composition and Temporal Fusion Based Inter Prediction for Learned Video Compression. IEEE Transactions on Circuits and Systems for Video Technology, 34(7):6460–6473, 2024. 1, 2, 3, 4, 6, 8
2024
-
[52]
Rethinking Alignment in Video Super-Resolution Transformers
Shuwei Shi, Jinjin Gu, Liangbin Xie, Xintao Wang, Yu- jiu Yang, and Chao Dong. Rethinking Alignment in Video Super-Resolution Transformers. Advances in Neural In- formation Processing Systems (NeurIPS) , 35:36081–36093,
-
[53]
Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network
Wenzhe Shi, Jose Caballero, Ferenc Husz ´ar, Johannes Totz, Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network. In Proceedings of the IEEE/CVF Conference on C...
-
[54]
Overview of the High Efficiency Video Coding (HEVC) Standard
Gary J Sullivan, Jens-Rainer Ohm, Woo-Jin Han, and Thomas Wiegand. Overview of the High Efficiency Video Coding (HEVC) Standard. IEEE Transactions on Circuits and Systems for Video Technology, 22(12):1649–1668, 2012. 1
2012
-
[55]
Standardized Exten- sions of High Efficiency Video Coding (HEVC).IEEE Jour- nal of selected topics in Signal Processing, 7(6):1001–1016,
Gary J Sullivan, Jill M Boyce, Ying Chen, Jens-Rainer Ohm, C Andrew Segall, and Anthony Vetro. Standardized Exten- sions of High Efficiency Video Coding (HEVC).IEEE Jour- nal of selected topics in Signal Processing, 7(6):1001–1016,
-
[56]
Offline and Online Optical Flow En- hancement for Deep Video Compression
Chuanbo Tang, Xihua Sheng, Zhuoyuan Li, Haotian Zhang, Li Li, and Dong Liu. Offline and Online Optical Flow En- hancement for Deep Video Compression. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 5118– 5126, 2024. 1
2024
-
[57]
RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
Zachary Teed and Jia Deng. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. In European Conference on Computer Vision (ECCV) , pages 402–419. Springer, 2020. 13
2020
-
[58]
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is All You Need. Advances in Neural Information Processing Systems (NeurIPS), 30, 2017. 3, 5
2017
-
[59]
MCL-JCV: a JND-based H.264/A VC video quality assessment dataset
Haiqiang Wang, Weihao Gan, Sudeng Hu, Joe Yuchieh Lin, Lina Jin, Longguang Song, Ping Wang, Ioannis Katsavouni- dis, Anne Aaron, and C-C Jay Kuo. MCL-JCV: a JND-based H.264/A VC video quality assessment dataset. In2016 IEEE International Conference on Image Processing (ICIP), ...
2016
-
[60]
Yao Wang and O. Lee. Use of two-dimensional deformable mesh structures for video coding .I. The synthesis problem: mesh-based function approximation and mapping. IEEE Transactions on Circuits and Systems for Video Technology, 6(6):636–646, 1996. 1
1996
-
[61]
Overview of the H.264/A VC Video Coding Standard
Thomas Wiegand, Gary J Sullivan, Gisle Bjontegaard, and Ajay Luthra. Overview of the H.264/A VC Video Coding Standard. IEEE Transactions on Circuits and Systems for Video Technology, 13(7):560–576, 2003. 1
2003
-
[62]
Affine Multipicture Motion-compensated Prediction
Thomas Wiegand, Eckehard Steinbach, and Bernd Girod. Affine Multipicture Motion-compensated Prediction. IEEE Transactions on Circuits and Systems for Video Technology, 15(2):197–209, 2005. 1
2005
-
[63]
Video Enhancement with Task-oriented Flow
Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and William T Freeman. Video Enhancement with Task-oriented Flow. International Journal of Computer Vision, 127:1106– 1125, 2019. 6
2019
-
[64]
Learned Low Bitrate Video Compres- sion with Space-time Super-resolution
Jiayu Yang, Chunhui Yang, Fei Xiong, Feng Wang, and Ronggang Wang. Learned Low Bitrate Video Compres- sion with Space-time Super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1786–1790, 2022. 3
2022
-
[65]
DNeRV: Model- ing Inherent Dynamics via Difference Neural Representation for Videos
Qi Zhao, M Salman Asif, and Zhan Ma. DNeRV: Model- ing Inherent Dynamics via Difference Neural Representation for Videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2031–2040, 2023. 2 11 Supplementary Material This supple...
2023
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.