REVIEW 4 major objections 5 minor 53 references
ViT-NeBLa: A Hybrid Vision Transformer and Neural Beer-Lambert Framework for Single-View 3D Reconstruction of Oral Anatomy from Panoramic Radiographs
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A hybrid transformer and X-ray-physics network reconstructs a 3D jaw volume from a single panoramic radiograph, outperforming prior dental reconstruction methods on synthetic test data.
desk verdict A sensible integration of known components with one genuinely useful sampling trick, but the missing NeBLa baseline and synthetic-only evaluation undercut the abstract's SOTA claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
$f_{\mathrm{fused}}(r,s) = f_{\mathrm{img}}(i(r),j(r)) + f_{\mathrm{pos}}(r,s)$ is the central fused feature vector, fed through an eight-layer MLP that predicts scalar densities from each query point. The image side comes from a hybrid ViT-CNN extractor: a Vision Transformer encoder-decoder with skip connections provides global context, and a small convolutional branch adds local texture. The position side comes from a learnable multi-resolution hash grid (16 levels, resolutions from 16 to 256, table size $2^{19}$, two features per level), which lifts each 3D coordinate into a compact higher-dimensional representation. Horseshoe-shaped focal-region sampling along tangent, non-intersecting elliptical rays cuts per-ray samples from 200 to 96 and eliminates intermediate density aggregation. A 3D U-Net refines the coarse rendered volume, and the composite loss combines voxel MSE, maximum-intensity projection consistency in three planes, and a VGG-16 perceptual term.
What would settle it
Run the trained model on real panoramic radiographs from patients who also have CBCT scans and compare the reconstructed volumes to the paired CBCT; if PSNR and SSIM fall to baseline levels or the jaw borders blur, the clinical claim fails. A faster check is to add focal-trough noise and ghost artifacts to the synthetic test images and measure how much PSNR and SSIM drop.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that the mapping from a single panoramic radiograph and a set of 3D query points to a volumetric density field can be learned end-to-end by fusing global ViT features, local CNN features, and multi-resolution hash position encodings into a compact MLP density predictor, then refining the rendered volume with a 3D U-Net. The geometric choice that makes the pipeline dependency-free is sampling 96 points per ray along non-intersecting rays tangent to an elliptical trajectory inside a horseshoe-shaped focal region, which removes the intermediate density aggregation step that intersecting-ray methods need. On a synthetic test set, the paper reports that this architecture outperforms 3DentAI, Oral-3D, and a residual CNN on PSNR, SSIM, and LPIPS, and that the full loss combination of voxel MSE, projection consistency, and perceptual features is what sharpens jaw borders and foramina.
Load-bearing premise
The central claim depends on the assumption that the computer-generated panoramic X-rays used for training and testing behave like real clinical panoramic X-rays; the paper trains only on synthetic images and acknowledges that real images carry more noise and ghost artifacts, so a large enough gap would void the clinical claim.
Editorial extensions
If this is right
- Single-view reconstruction no longer needs CBCT flattening or dental arch curves, removing two inputs that are often unavailable in routine clinical settings.
- The horseshoe sampling cuts per-ray sample points by 52 percent (96 versus 200), roughly halving memory and compute while preserving reported fidelity.
- The hybrid ViT-CNN extractor and hash encoding together produce sharper cortical borders, gonial angle, ramus, and foramina than U-Net-style baselines on synthetic inputs.
- Training on a CBCT cohort that includes metallic implants and edentulous segments yields a model that keeps its quantitative advantage across diverse anatomy and artifacts.
- The projection-consistency and perceptual losses are necessary for recovering fine trabecular detail; MSE alone leaves the volume overly smooth.
Reading between the lines
- If the synthetic-to-real gap is closed with image translation (the paper names this as future work), the same architecture could plausibly reconstruct bone as well as teeth from routine clinical PX, going beyond tooth-centric reconstructions.
- The tangent-elliptical, non-intersecting ray formulation removes per-patient arch curves and may transfer to panoramic radiographs from other scanners, since the trajectory is defined by the scan geometry rather than by patient-specific priors.
- A small set of real PX-CBCT pairs used for fine-tuning or evaluation would directly test the clinical boundary; the paper's reported performance drop with noisier real PX currently marks an untested limit.
- Because the output is an implicit density field rather than an explicit surface, downstream surgical planning would need a mesh extraction step; the paper notes this and points to Gaussian splatting as a future explicit representation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes ViT-NeBLa, a hybrid Vision Transformer and Neural Beer-Lambert framework for reconstructing a 3D CBCT-like volume from a single panoramic radiograph (PX). The method combines a hybrid ViT-CNN feature extractor, learnable multi-resolution hash positional encoding, an 8-layer MLP density predictor, and a 3D U-Net refinement module, trained with MSE, projection-consistency, and perceptual losses. Synthetic PX images are generated from 600 CBCT volumes by ray tracing along an elliptical trajectory within a horseshoe-shaped focal region, sampling 96 points per ray versus 200 in the prior NeBLa method. On a held-out synthetic test set, the paper reports PSNR 23.48 dB, SSIM 74.93%, and LPIPS 0.4204, outperforming three baselines (residual CNN, Oral-3D autoencoder, 3DentAI). The paper claims state-of-the-art performance and positions the method as a cost-effective, radiation-efficient clinical alternative.
Significance. The engineering components are described in enough detail to be reproduced in principle: configuration tables (Tables V-VI), the loss formulation (Eqs. 14-18), and the synthetic data-generation pipeline (Appendix A) are explicit, and the loss ablation in Table IV isolates the contribution of the perceptual term. The reported gains over the three chosen baselines are internally consistent, and the 52% reduction in per-ray sampling is a concrete efficiency claim. However, the significance as stated is undercut by two omissions: the direct predecessor NeBLa is absent from all quantitative comparisons, and every evaluation uses synthetic PX generated by the authors' own forward model. The clinical claim in the abstract therefore goes beyond the evidence presented in the manuscript.
major comments (4)
- [Section IV-C, Tables II and III] NeBLa [9], the method the paper extends by name and the only approach whose sampling strategy is directly contrasted in Section III and Fig. 2, is missing from the quantitative comparison. The abstract's claim of significantly outperforming prior state-of-the-art methods cannot be assessed without reporting NeBLa under the same training and test protocol; a re-implementation or a comparable published result on the same data is required.
- [Section IV-A1b and Section V] All training and test images are synthetic PX produced by ray tracing the same CBCT volumes (600 scans, 8:1:1 split), and the authors state that traditional PX images retain unwanted features such as ghost artifacts, resulting in higher noise and reduced model performance when used as input. The abstract's clinical conclusion, that the method offers a cost-effective, radiation-efficient alternative for enhanced dental diagnostics, is therefore not supported by any real-PX evaluation; either real panoramic images must be tested or the claims must be restricted to synthetic PX.
- [Appendix A and Eqs. (3)-(4)] The query point set P is defined by the elliptical trajectory and horseshoe focal region extracted from the CBCT volume during synthetic generation. At inference on a real PX, this trajectory and focal region are unknown and would need to be estimated from the 2D image alone; the manuscript reports no such estimation procedure and no sensitivity analysis with respect to trajectory parameters. This is a load-bearing gap for the proposed use of the method on clinical PX.
- [Section IV-D, Table IV] The ablation study varies only the loss terms. The three claimed architectural innovations, namely the ViT-CNN hybrid versus a U-Net/CNN extractor, hash positional encoding versus Fourier encoding, and 96-point horseshoe sampling versus NeBLa's 200-point sampling, are not ablated, so the contribution of each to the reported PSNR/SSIM/LPIPS gains is unsubstantiated.
minor comments (5)
- [Abstract and Section VI] The abstract and conclusion repeatedly say 'single PX' while Section IV-A1b clarifies that the experiments use synthetic PX; please make this distinction consistent throughout the paper.
- [Section IV-A1c] The paragraph lists 'clipping, log-compression, standardization, and min-max rescaling,' but the preceding description mentions clipping, Z-score normalization, and linear rescaling with no log-compression step; please reconcile the description.
- [Table II caption and Section III] There are typographical errors that should be corrected: 'Quantative' in the Table II caption, 'anatomoy' in the introduction of Section III, and 'metrices' in Section IV-A2.
- [Appendix B, Eq. (14)] The loss weights are stated as λ1=1/1.2 and λ2=1/25 only in Appendix B, with no sensitivity study; please state how these values were chosen and whether results are stable to them.
- [Figure 2 caption] The caption says both methods use identical inter-sample spacing, while Eq. (4) defines uniform sampling with S=96 over the focal interval; please clarify whether this refers to equal spacing in the parameterization or in physical space.
Circularity Check
No circularity: the reconstruction benchmark is a supervised held-out inverse problem, and the synthetic-data realism gap is a stated limitation rather than a circular derivation.
full rationale
The paper's derivation chain is a supervised learning pipeline: synthetic panoramic radiographs are generated from CBCT volumes by Beer-Lambert ray casting, and the model is trained to invert that forward model, mapping a held-out 2D projection plus 3D query coordinates to the corresponding CBCT volume. The training/validation/test split (8:1:1 of 600 patient cases) ensures that the reported PSNR, SSIM, and LPIPS values are computed on data the network never saw during training; no parameter is fitted to the test set, and none of the loss terms is defined in terms of the reported metrics. The query rays and horseshoe-shaped focal region are determined by the same geometric forward model that creates the synthetic PX, but this is the intended inverse-task structure rather than a definitional identity: a single projection is ill-posed, and the network must learn a prior to select a plausible volume. The citation to the authors' prior work [48] for dataset preparation is a self-citation, but it is not load-bearing: the geometry is reproduced in Appendix A and rests on the standard Beer-Lambert attenuation model, not on a uniqueness theorem or an unstated ansatz. The paper's Limitations section explicitly concedes that traditional PX images contain noise and artifacts not present in synthetic data, causing reduced performance; that is an honest generalization limitation, not a circular step. Thus the core reconstruction claim is self-contained with respect to the stated synthetic benchmark, and the clinical real-PX claim remains an external validity risk rather than a circularity.
Assumptions & free parameters
free parameters (5)
- Loss weights lambda1, lambda2 =
lambda1 = 1/1.2, lambda2 = 1/25
- Samples per ray S =
96
- Hash grid hyperparameters =
xi=16 levels, r0=16, rmax=256, F=2, table size 2^19
- Network architecture sizes =
MLP depth 8 width 128, U-Net channels 64/128/256/512, ViT 12 layers, patch 16x16
- Elliptical trajectory geometry =
patient-specific, derived from jaw contour
assumptions (5)
- domain assumption Beer-Lambert law forward projection models X-ray attenuation
- domain assumption Synthetic PX images are representative of clinical panoramic radiographs
- domain assumption Tangent rays to the elliptical trajectory do not intersect inside the focal region
- domain assumption A neural network can learn the PX-to-CBCT mapping from paired synthetic data
- standard math NeRF-style volume rendering approximates CBCT density integration
Cite this review
Pith. "Pith review of ViT-NeBLa: A Hybrid Vision Transformer and Neural Beer-Lambert Framework for Single-View 3D Reconstruction of Oral Anatomy from Panoramic Radiographs." pith.science (2026). https://pith.science/paper/I3LBNN3U
@misc{pith2026250613195,
author = {Pith},
title = {Pith review of: ViT-NeBLa: A Hybrid Vision Transformer and Neural Beer-Lambert Framework for Single-View 3D Reconstruction of Oral Anatomy from Panoramic Radiographs},
year = {2026},
howpublished = {\url{https://pith.science/paper/I3LBNN3U}},
note = {Machine review of arXiv:2506.13195}
}
abstract
Dental diagnosis relies on two primary imaging modalities: panoramic radiographs (PX) providing 2D oral cavity representations, and Cone-Beam Computed Tomography (CBCT) offering detailed 3D anatomical information. While PX images are cost-effective and accessible, their lack of depth information limits diagnostic accuracy. CBCT addresses this but presents drawbacks including higher costs, increased radiation exposure, and limited accessibility. Existing reconstruction models further complicate the process by requiring CBCT flattening or prior dental arch information, often unavailable clinically. We introduce ViT-NeBLa, a vision transformer-based Neural Beer-Lambert model enabling accurate 3D reconstruction directly from single PX. Our key innovations include: (1) enhancing the NeBLa framework with Vision Transformers for improved reconstruction capabilities without requiring CBCT flattening or prior dental arch information, (2) implementing a novel horseshoe-shaped point sampling strategy with non-intersecting rays that eliminates intermediate density aggregation required by existing models due to intersecting rays, reducing sampling point computations by $52 \%$, (3) replacing CNN-based U-Net with a hybrid ViT-CNN architecture for superior global and local feature extraction, and (4) implementing learnable hash positional encoding for better higher-dimensional representation of 3D sample points compared to existing Fourier-based dense positional encoding. Experiments demonstrate that ViT-NeBLa significantly outperforms prior state-of-the-art methods both quantitatively and qualitatively, offering a cost-effective, radiation-efficient alternative for enhanced dental diagnostics.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[9]
J. B. Ludlow, R. Timothy, C. Walker, R. Hunter, E. Be- navides, D. B. Samuelson, and M. J. Scheske, Dentomax- illofacial Radiology 44, 20140197 (2015)
work page 2015
-
[1]
Implementation details a. Dataset collection. The study used CBCT clinical head scans 623, comprising 408 women and 215 men with an average age of 50 years and a standard de- viation of 26 years. These scans were obtained from the Chosun School of Dentistry in Gwangju, South Korea. Image acquisition was performed with the Carestream Health CS900 scanner a...
-
[2]
Evaluation metrics To assess image reconstruction results, the 2D-to-3D reconstruction model was evaluated with the standard metrices as follows: • PSNR : Peak Signal-to-Noise Ratio (PSNR) mea- sures the quality of a reconstructed image by com- paring it to the original, focusing on the ratio be- tween the maximum possible signal power and the power of co...
-
[3]
W. C. Scarfe, A. G. Farman, and P. Sukovic, Journal (Canadian Dental Association) 72, 75—80 (2006)
work page 2006
-
[4]
C. Angelopoulos, W. C. Scarfe, and A. G. Farman, At- las of the Oral and Maxillofacial Surgery Clinics 20, 1 (2012)
work page 2012
- [5]
-
[6]
S. D. Ganz, Dental Clinics of North America 55, 515 (2011)
work page 2011
- [7]
Show all 53 references
-
[8]
Gupta and S
J. Gupta and S. Ali, National Journal of Maxillofacial Surgery 4, 2 (2013)
2013
-
[10]
A. C. Oenning, R. Jacobs, R. Pauwels, A. Stratis, M. Hedesiu, B. Salmon, and h. On behalf of the DIM- ITRA Research Group, Pediatric Radiology 48, 308 (2018)
2018
-
[11]
S. Park, S. Kim, D. Kwon, Y. Jang, I.-S. Song, and S. J. Baek, Proceedings of the AAAI Conference on Artificial Intelligence 38, 4433 (2024)
2024
-
[12]
Vandenberghe, R
B. Vandenberghe, R. Jacobs, and H. Bosmans, European Radiology 20, 2637 (2010)
2010
-
[13]
Shah, World Journal of Radiology 6, 794 (2014)
N. Shah, World Journal of Radiology 6, 794 (2014)
2014
-
[14]
A. P. S and W. You, in2023 International Technical Con- ference on Circuits/Systems, Computers, and Communi- cations (ITC-CSCC) (IEEE, Jeju, Korea, Republic of,
-
[16]
Stramotas, The European Journal of Orthodontics 24, 43 (2002)
S. Stramotas, The European Journal of Orthodontics 24, 43 (2002)
2002
-
[17]
B. F. Gribel, M. N. Gribel, D. C. Fraz˜ ao, J. A. McNa- mara, and F. R. Manzi, The Angle Orthodontist 81, 26 (2011)
2011
-
[18]
Litjens, T
G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Se- tio, F. Ciompi, M. Ghafoorian, J. A. Van Der Laak, B. Van Ginneken, and C. I. S´ anchez, Medical Image Anal- ysis 42, 60 (2017)
2017
-
[19]
D. Shen, G. Wu, and H.-I. Suk, Annual Review of Biomedical Engineering 19, 221 (2017)
2017
-
[20]
Y. Bai, L. Wong, and T. Twan, Survey on fundamen- tal deep learning 3d reconstruction techniques (2024), arXiv:2407.08137 [cs.CV]
2024 arXiv
-
[21]
Esteva, A
A. Esteva, A. Robicquet, B. Ramsundar, V. Kuleshov, M. DePristo, K. Chou, C. Cui, G. Corrado, S. Thrun, and J. Dean, Nature Medicine 25, 24 (2019)
2019
-
[22]
C. Chen, C. Qin, H. Qiu, G. Tarroni, J. Duan, W. Bai, and D. Rueckert, Frontiers in Cardiovascular Medicine7, 25 (2020)
2020
-
[23]
W. Song, Y. Liang, J. Yang, K. Wang, and L. He, Pro- ceedings of the AAAI Conference on Artificial Intelli- gence 35, 566 (2021)
2021
-
[24]
Yaqub, F
M. Yaqub, F. Jinchao, K. Arshid, S. Ahmed, W. Zhang, M. Z. Nawaz, and T. Mahmood, Computational and Mathematical Methods in Medicine 2022, 1 (2022)
2022
-
[25]
R. Zha, Y. Zhang, and H. Li, in Medical Image Com- puting and Computer Assisted Intervention – MICCAI 2022, edited by L. Wang, Q. Dou, P. T. Fletcher, S. Spei- del, and S. Li (Springer Nature Switzerland, Cham, 2022) pp. 442–452
2022
-
[26]
W. Song, H. Zheng, D. Tu, C. Liang, and L. He, Oral-3dv2: 3d oral reconstruction from panoramic x- ray imaging with implicit neural representation (2023), arXiv:2303.12123 [eess.IV]
2023 arXiv
-
[27]
Jader, J
G. Jader, J. Fontineli, M. Ruiz, K. Abdalla, M. Pithon, and L. Oliveira, in 2018 31st SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI) (IEEE, Parana, 2018) pp. 400–407
2018
-
[28]
Lee, D.-h
J.-H. Lee, D.-h. Kim, S.-N. Jeong, and S.-H. Choi, Jour- nal of Periodontal & Implant Science 48, 114 (2018)
2018
-
[29]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (2021), arXiv:2010.11929
2021 arXiv
-
[30]
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, ACM Computing Surveys 54, 1 (2022)
2022
-
[31]
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, in 2021 IEEE/CVF International Confer- ence on Computer Vision (ICCV) (IEEE, Montreal, QC, Canada, 2021) pp. 9992–10002
2021
-
[32]
J. Chen, Y. He, E. C. Frey, Y. Li, and Y. Du, Vit-v-net: Vision transformer for unsupervised volumetric medical image registration (2021), arXiv:2104.06468 [eess.IV]
2021 arXiv
-
[33]
K.-E. Lin, L. Yen-Chen, W.-S. Lai, T.-Y. Lin, Y.-C. Shih, and R. Ramamoorthi, Vision transformer for nerf- based view synthesis from a single input image (2022), arXiv:2207.05736 [cs.CV]
2022 arXiv
-
[34]
Mildenhall, P
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Bar- ron, R. Ramamoorthi, and R. Ng, Communications of the ACM 65, 99 (2022)
2022
-
[35]
X. Wang, Y. Chen, S. Hu, H. Fan, H. Zhu, and X. Li, Neural radiance fields in medical imaging: A survey (2025), arXiv:2402.17797 [eess.IV]
2025 arXiv
-
[36]
C. Zhu, R. Ishikawa, M. Kagesawa, T. Yuzawa, T. Wat- suji, and T. Oishi, Neas: 3d reconstruction from x- ray images using neural attenuation surface (2025), arXiv:2503.07491 [eess.IV]
2025 arXiv
-
[37]
Y. Yang, Z. Cui, C. Li, and W. Wang, Toothinpaintor: Tooth inpainting from partial 3d dental model and 2d panoramic image (2022), arXiv:2211.15502 [cs.CV]
2022 arXiv
-
[38]
Y. Lin, Z. Luo, W. Zhao, and X. Li, in Medical Image Computing and Computer Assisted Intervention – MIC- CAI 2023 , Vol. 14229, edited by H. Greenspan, A. Mad- abhushi, P. Mousavi, S. Salcudean, J. Duncan, T. Syeda- Mahmood, and R. Taylor (Springer Nature Switzerland, Cham, 20...
2023
-
[39]
Molaei, A
A. Molaei, A. Aminimehr, A. Tavakoli, A. Kazerouni, B. Azad, R. Azad, and D. Merhof, in 2023 IEEE/CVF International Conference on Computer Vision Work- shops (ICCVW) (IEEE, Paris, France, 2023) pp. 2373– 2383. 19
2023
-
[40]
Y. Fang, L. Mei, C. Li, Y. Liu, W. Wang, Z. Cui, and D. Shen, Snaf: Sparse-view cbct reconstruction with neu- ral attenuation fields (2022), arXiv:2211.17048 [eess.IV]
2022 arXiv
-
[41]
S. Park, S. Kim, I.-S. Song, and S. J. Baek, in Medical Image Computing and Computer Assisted Intervention - MICCAI 2023 , Vol. 14229, edited by H. Greenspan, A. Madabhushi, P. Mousavi, S. Salcudean, J. Duncan, T. Syeda-Mahmood, and R. Taylor (Springer Nature Switzerland, Cham...
2023
-
[42]
P.Sunilkumar, S
A. P.Sunilkumar, S. Moon, and W. Yoo, Journal of the Information Processing Society 13, 326 (2024)
2024
-
[43]
X. Li, M. Meng, Z. Huang, L. Bi, E. Delamare, D. Feng, B. Sheng, and J. Kim, in Medical Image Computing and Computer Assisted Intervention – MICCAI 2024 , Vol. 15007, edited by M. G. Linguraru, Q. Dou, A. Feragen, S. Giannarou, B. Glocker, K. Lekadir, and J. A. Schnabel (Sprin...
2024
-
[44]
Liang, W
Y. Liang, W. Song, J. Yang, L. Qiu, K. Wang, and L. He, in Medical Image Computing and Computer As- sisted Intervention – MICCAI 2020 , Vol. 12262, edited by A. L. Martel, P. Abolmaesumi, D. Stoyanov, D. Ma- teus, M. A. Zuluaga, S. K. Zhou, D. Racoceanu, and L. Joskowicz (Spri...
2020
-
[45]
W. Ma, H. Wu, Z. Xiao, Y. Feng, J. Wu, and Z. Liu, in Medical Image Computing and Computer Assisted Inter- vention – MICCAI 2024 , Vol. 15003, edited by M. G. Lin- guraru, Q. Dou, A. Feragen, S. Giannarou, B. Glocker, K. Lekadir, and J. A. Schnabel (Springer Nature Switzer- la...
2024
-
[46]
R. A. Ketcham and R. D. Hanna, Computers & Geo- sciences 67, 49 (2014)
2014
-
[47]
Max, IEEE Transactions on Visualization and Com- puter Graphics 1, 99 (1995)
N. Max, IEEE Transactions on Visualization and Com- puter Graphics 1, 99 (1995)
1995
-
[48]
X.-Y. Zhou, P. Li, Z.-Y. Wang, and G.-Z. Yang, in Mul- tiscale Multimodal Medical Imaging , edited by Q. Li, R. Leahy, B. Dong, and X. Li (Springer International Publishing, Cham, 2020) pp. 101–108
2020
-
[49]
Ramachandran, B
P. Ramachandran, B. Zoph, and Q. V. Le, Searching for activation functions (2017), arXiv:1710.05941
2017 arXiv
-
[50]
A. P. Sunilkumar, B. K. Parida, S.-Y. Moon, and W. You, IEEE Access , 1 (2025)
2025
-
[51]
Z. Wang, E. Simoncelli, and A. Bovik, in The Thrity- Seventh Asilomar Conference on Signals, Systems & Computers, 2003 (IEEE, Pacific Grove, CA, USA, 2003) pp. 1398–1402
2003
-
[52]
Zhang, P
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (IEEE, Salt Lake City, UT, 2018) pp. 586–595
2018
-
[53]
Henzler, V
P. Henzler, V. Rasche, T. Ropinski, and T. Ritschel, Computer Graphics Forum 37, 377 (2018)
2018
-
[54]
Szymanowicz, C
S. Szymanowicz, C. Rupprecht, and A. Vedaldi, in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, Seattle, WA, USA, 2024) pp. 10208–10217
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.