REVIEW 3 major objections 5 minor 46 references
AcoustiTrace: When Plausible Sound Violates Physics
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read AcoustiTrace shows that plausible, synchronized sound in generated video does not guarantee faithful modeling of the tested acoustic relations.
desk verdict A serious, unusually transparent benchmark for acoustic physical realism in A/V generation; the RT60 visual estimator is the main soft spot and the refinement study is only a proof of concept. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the relation-level diagnostic contract written as $D_k=(V_k,A_k,R_k,G_k,S_k)$. For each dimension it extracts required visual evidence, measures the matched audio quantity, states the expected acoustic relation, checks validity, and maps residual to a score. Example relations include the inverse-square law for range attenuation, Sabine's formula $T_{60}=0.161V/A$ for RT60 consistency, and exponential energy decay for impact decay. The same residual used for scoring can be turned into a differentiable guidance objective.
What would settle it
Measure the visual RT60 estimator's predictions on a held-out set of rooms never seen in training, using independently measured reverberation times rather than simulated labels; if the visual estimates disagree with the measured values by more than the benchmark's tolerance for many cases, the RT60-based failure rankings would not stand.
Extended reading notes
Core claim
The paper's central discovery is that perceptual plausibility and synchronization are not proxies for acoustic physical fidelity. When nine joint audio-video generators are evaluated on eight relation-level tests, no model consistently satisfies the expected relations; scores are frequently near chance for Log Attack Time, RT60 Consistency, and Range Attenuation, while Causality Violation and local event relations are generally high. The paper states this as: 'plausible and synchronized sound events do not guarantee faithful modeling of the tested acoustic relations.' It further shows that the diagnosed range-attenuation residual can be used as a differentiable guidance objective, improving the targeted relation in 80.16% of valid samples without retraining.
Load-bearing premise
The RT60 Consistency dimension assumes that a visual estimator trained on simulated room labels from the same room set it is tested on accurately estimates apparent reverberation in real and generated scenes, with supporting real-world evidence limited to 26 proxy pairs and one illustrative measured-room case.
Editorial extensions
If this is right
- No evaluated generator dominates across all eight dimensions; leadership is split, implying the models are not uniformly physical.
- Relations that are locally explicit in training clips (causality, lateral stability, motion-loudness, impact decay) score higher than relations requiring longer geometric or environmental consistency (log attack time, RT60 consistency, range attenuation).
- A diagnosed residual can guide inference: range-guided audio resampling raises mean $R^2$ from 0.714 to 0.858, improves text-audio alignment, and preserves loudness, without retraining.
- Embedding-transition metrics and relation-specific acoustic diagnostics are complementary: a high embedding-based score can coexist with a low attenuation score and vice versa.
Reading between the lines
- The paper leaves implicit that the same diagnostic contract could be reused for reward modeling in reinforcement-learning fine-tuning, because a per-output relation residual is exactly a reward signal.
- If the range-attenuation result generalizes, other continuous acoustic relations such as approach gain and lateral stability may also be correctable by gradient guidance on the decoded audio envelope, without retraining.
- A scene-disjoint validation of the visual RT60 estimator would separate estimator error from generator error; this is a natural next experiment the paper does not run.
- The weak correlation with embedding-transition scores suggests future benchmarks should report both types of evidence, since each can miss what the other catches.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. AcoustiTrace is a diagnostic benchmark for acoustic physical realism in joint text-to-audio-video (T2AV) and image-to-audio-video (I2AV) generation. It defines eight evaluation dimensions spanning sound generation, propagation environment, and acoustic reception; constructs a dataset of 11,296 real-world audio-video clips and 82,828 RGB-D observations with simulated 500 Hz RT60 labels; builds targeted prompt suites (605 T2AV and 748 I2AV prompts); validates evaluators through real-world relation recovery, controlled perturbations, and human agreement; and evaluates nine joint audio-video generators. The authors report that plausible, synchronized sound does not guarantee faithful modeling of acoustic relations, with RT60 Consistency, Range Attenuation, and Log Attack Time being particularly challenging. A final intervention converts the diagnosed range-attenuation residual into a differentiable audio-sampling objective, improving the targeted metric in 80.16% of valid samples while largely preserving non-target quality metrics.
Significance. If the identified limitations are addressed, AcoustiTrace would be a valuable community benchmark: it organizes evaluation around interpretable acoustic mechanisms, explicitly gates validity, and validates evaluators with controlled perturbations, human judgments, bootstrap confidence intervals, and matched-valid analyses. These strengths are genuinely above the norm for audio-video generation benchmarks. The headline conclusion is credible for several dimensions, but the RT60 Consistency evidence is weakened by limited external validation of the visual estimator, and the intervention is more an optimization feasibility study than a demonstrated general refinement method. With a re-scoped RT60 dimension and a softened intervention claim, the paper could become a reference point for physics-aware evaluation of joint audio-video models.
major comments (3)
- [S2.2, S6.3, S6.4, Table 1] The visual RT60 estimator is trained and evaluated on SoundSpaces 2.0 simulated labels from the same 83 Matterport3D scenes with a non-scene-disjoint split, and its real-world support is limited to 26 valid STARSS23 proxy pairs and one illustrative BRAS CR3 case. The paper itself states in S2.2 that the estimator is 'not presented as a measurement-grade estimator for arbitrary real rooms.' Since RT60 Consistency is one of the headline, lowest scores in Table 1 (e.g., Seedance 2.0 T2AV 24.28 and Wan 2.7 T2AV 33.20), the current evidence cannot rule out that these results partly reflect estimator distribution shift on generated videos rather than genuine acoustic violations. Please re-validate with a scene-disjoint split and substantially more measured-room evidence, or explicitly demote the RT60 dimension to an exploratory sub-diagnostic and remove it from the paper's central claims.
- [S5.2, Tables S6, S10, S11] Table 1 reports conditional means over valid outputs, and RT60 Consistency has the lowest and most variable validity rates (40.5% to 82.1%). The matched-valid analysis covers only three models per task, and within that analysis RT60 is the only dimension whose ordering changes (JavisDiT++ and Ovi swap in T2AV). No matched-valid evidence is provided for Seedance 2.0 and Wan 2.7, the two models with the lowest T2AV RT60 scores. The cross-model RT60 rankings in Table 1 are therefore not robust to validity-coverage differences and should be either fully matched-validated across all reported models or excluded from the headline comparisons.
- [S8.1, Eq. (S17), Table 2] The range-guided intervention optimizes a loss that is a differentiable surrogate of the Range Attenuation evaluator's own residual (log audio envelope versus log visual range). Improving mean R2 from 0.7138 to 0.8580 and winning in 80.16% of samples is therefore expected if the optimization is successful; it does not by itself demonstrate that AcoustiTrace diagnostics transfer to 'model refinement' beyond optimizing the same measurement. The non-target metrics (CLAP, LUFS, PQ) show the effect is not a global loudness adjustment, and the CLAP improvement is encouraging, but the authors should either add held-out prompts/models or human perceptual judgments, or explicitly reframe the study as a feasibility demonstration of converting a diagnostic residual into an optimization signal.
minor comments (5)
- [Table 1, S5.2] The term 'conditional scores' is used in the main text without definition; please state explicitly that all scores are means over valid outputs, with validity rates reported separately.
- [S4] The unique T2AV prompt-count formula (84+222+110+111+84-6=605) is confusing because the text lists 222 each for Approach Gain and Lateral Stability and separately lists Causality Violation; clarify that Approach and Lateral share a common receiver-motion pool and that Causality Violation reuses prompts from other pools.
- [S6.5] The human evaluation reports agreement between evaluator preferences and human judgments, but not inter-rater agreement; please report a chance-corrected agreement statistic such as Fleiss' kappa to show that raters agree with each other.
- [S3.2, Eq. (S10)] The RT60 consistency score saturates for all ratios within a factor of 1.5, which can compress meaningful differences between generators; please state this saturation behavior and its implication for cross-model comparisons in the main text.
- [General] The paper should include an artifact availability statement in the main text, since the reproducibility of a benchmark depends on public release of code, evaluator weights, and the data manifest.
Circularity Check
Benchmark evaluation is not circular; only the range-guidance intervention improves its target by construction.
-
self definitional
[Main paper, 'Range-Guided Audio Sampling'; Supplementary Eq. S17 and S8.1]
"We target the identified Range Attenuation failure and formulate the expected attenuation relation used by the evaluator as a differentiable guidance objective for audio sampling. ... Distance guidance increases the mean Range Attenuation R2 from 0.7138 to 0.8580 and outperforms unguided resampling in 80.16% of cases."
The guidance loss in Eq. S17 is a log-ratio envelope objective with gamma=1, i.e., log(e_i/e_j) = log(d_j/d_i), which is exactly the inverse-distance amplitude relation underlying the Range Attenuation evaluator (Eq. S12: Delta L* = -20 log10(d(t)/d(t0))). Optimizing this objective on the fixed video directly maximizes the evaluator's R2, so the reported 80.16% improvement is a manipulation check rather than an independent discovery. The paper's non-target metrics (CLAP, LUFS, PQ) are independent and show the intervention is not wholly vacuous, but the headline 'improvement in the targeted relation' reduces by construction to optimizing the measured relation.
full rationale
AcoustiTrace's benchmark evaluation is not circular: each of the eight dimensions tests an externally specified acoustic relation (inverse-square attenuation, Sabine reverberation, exponential decay, causality) against independently measured visual and audio evidence, and the evaluators are validated with controlled perturbations and human agreement. The RT60 visual-estimator limitation, including the non-scene-disjoint train/test split and sparse real-world proxy validation, is a validity and robustness concern rather than a circularity: the estimator is honestly scoped and not used to define the physical relation. No load-bearing self-citation chain is present. The only construction-redundant element is the intervention study, where Eq. S17 optimizes the same log-ratio range-loudness relation that the Range Attenuation evaluator scores, making the reported R2 and 80.16% win expected; the independent CLAP, LUFS, and Audiobox PQ metrics provide partial independent content. Overall, the central benchmark derivation is self-contained, so the score is moderate rather than severe.
Assumptions & free parameters
free parameters (6)
- RT60 consistency tolerance factor =
score max at ratio 1.5, zero at ratio 3.0
- Log Attack Time exponential scale =
0.35 in Equation S3
- Causality violation scoring margin =
1 ms
- Impact decay fit window and tail penalty =
fit from 0.02 to 0.50 s after peak, tail residual normalized by dynamic range
- Range attenuation sliding-window configuration =
0.40 s windows, 0.05 s stride; minimum 1.50 s dominant segment; relative-depth thresholds
- Range-guided sampling gamma exponent =
gamma = 1 (inverse-distance amplitude attenuation)
assumptions (7)
- domain assumption Ideal free-field inverse-square law applies to range attenuation with approximately constant source power, stable directivity, and negligible reflections.
- domain assumption Sabine's diffuse-field relation T60 = 0.161 V / A approximates real room reverberation.
- domain assumption Simulated SoundSpaces 2.0 room impulse responses at 500 Hz provide valid RT60 supervision for real and generated audio.
- domain assumption Semantic-material matching of PTB absorption coefficients to RGB-D segmentation yields surface absorption maps accurate enough for visual RT60 estimation.
- domain assumption Apparent RT60 estimated from generated audio via a Schroeder decay curve is comparable to the room RT60 implied by visual cues.
- domain assumption The decoded-mel envelope extracted from the Audio VAE is an amplitude-like proxy whose logarithm scales linearly with distance attenuation.
- domain assumption Two-reviewer screening and automated MLLM/detector pipelines reliably identify when a target acoustic relation is observable and measurable.
Cite this review
Pith. "Pith review of AcoustiTrace: When Plausible Sound Violates Physics." pith.science (2026). https://pith.science/paper/EM26T47E
@misc{pith2026260802035,
author = {Pith},
title = {Pith review of: AcoustiTrace: When Plausible Sound Violates Physics},
year = {2026},
howpublished = {\url{https://pith.science/paper/EM26T47E}},
note = {Machine review of arXiv:2608.02035}
}
read the original abstract
Recent audio-video generators can produce semantically plausible and apparently synchronized sound, yet may still violate the acoustic processes implied by visible events and environments. Existing benchmarks provide limited support for attributing such violations to particular acoustic processes and quantifying their severity. We introduce AcoustiTrace, a diagnostic benchmark that formalizes acoustic physical realism in audio-video generation. AcoustiTrace organizes text-to-audio-video (T2AV) and image-to-audio-video (I2AV) evaluation around the acoustic process, covering sound generation, propagation environment, and acoustic reception through eight dimensions grounded in measurable acoustic quantities. Based on these evaluation dimensions, we construct a large-scale dataset organized around acoustic mechanisms, comprising real-world audio-video recordings and acoustically annotated RGB-D observations, and use it to develop targeted prompt suites and validated evaluators. Experiments reveal that even leading generators still struggle with fundamental acoustic processes despite producing plausible sound events. Finally, we show that the diagnostics AcoustiTrace provides for specific acoustic relations can guide model refinement toward more physically faithful audio and open new directions for incorporating acoustic principles into training objectives, reward modeling, and candidate selection.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
A Benchmark for Room Acoustical Simulation: Concept and Database , journal =
Brinkmann, Fabian and Asp. A Benchmark for Room Acoustical Simulation: Concept and Database , journal =. 2021 , doi =
work page 2021
-
[2]
Asp. 2020 , howpublished =. doi:10.14279/depositonce-6726.3 , note =
-
[3]
Do Joint Audio-Video Generation Models Understand Physics? , author =. 2026 , eprint =
work page 2026
-
[4]
2023 IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation , author =. 2023 IEEE International Conference on Acoustics, Speech and Signal Processing , pages =. 2023 , doi =
work page 2023
-
[5]
Girdhar, Rohit and El-Nouby, Alaaeldin and Liu, Zhuang and Singh, Mannat and Alwala, Kalyan Vasudev and Joulin, Armand and Misra, Ishan , booktitle =
-
[6]
Iashin, Vladimir and Xie, Weidi and Rahtu, Esa and Zisserman, Andrew , booktitle =. 2024 , doi =
work page 2024
-
[7]
Zhang, Yiming and Gu, Yicheng and Zeng, Yanhong and Xing, Zhening and Wang, Yuancheng and Wu, Zhizheng and Liu, Bin and Chen, Kai , journal =. 2026 , doi =
work page 2026
-
[8]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Read, Watch and Scream! Sound Generation from Text and Video , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2025 , doi =
work page 2025
Show all 46 references
-
[9]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2024 , doi =
2024
-
[10]
Cheng, Ho Kei and Ishii, Masato and Hayakawa, Akio and Shibuya, Takashi and Schwing, Alexander and Mitsufuji, Yuki , booktitle =
-
[11]
2508.16930 , archivePrefix =
Shan, Sizhe and Li, Qiulin and Cui, Yutao and Yang, Miles and Wang, Yuehai and Yang, Qun and Zhou, Jin and Zhong, Zhao , year =. 2508.16930 , archivePrefix =
-
[12]
2506.21448 , archivePrefix =
Liu, Huadai and Luo, Kaicheng and Wang, Jialei and Wang, Wen and Chen, Qian and Zhao, Zhou and Xue, Wei , year =. 2506.21448 , archivePrefix =
-
[13]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Seeing and Hearing: Open-Domain Visual-Audio Generation with Diffusion Latent Aligners , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[14]
2406.07686 , archivePrefix =
Wang, Kai and Deng, Shijian and Shi, Jing and Hatzinakos, Dimitrios and Tian, Yapeng , year =. 2406.07686 , archivePrefix =
-
[15]
and Shi, Yangyang and Chandra, Vikas , year =
Liu, Haohe and Le Lan, Gael and Mei, Xinhao and Ni, Zhaoheng and Kumar, Anurag and Nagaraja, Varun and Wang, Wenwu and Plumbley, Mark D. and Shi, Yangyang and Chandra, Vikas , year =. 2412.15220 , archivePrefix =
-
[16]
2026 , url =
Liu, Kai and Li, Wei and Chen, Lai and Wu, Shengqiong and Zheng, Yanhao and Ji, Jiayi and Zhou, Fan and Luo, Jiebo and Liu, Ziwei and Fei, Hao and Chua, Tat-Seng , booktitle =. 2026 , url =
2026
-
[17]
2509.06155 , archivePrefix =
Wang, Duomin and Zuo, Wei and Li, Aojie and Chen, Ling-Hao and Liao, Xinyao and Zhou, Deyu and Yin, Zixin and Dai, Xili and Jiang, Daxin and Yu, Gang , year =. 2509.06155 , archivePrefix =
-
[18]
2510.01284 , archivePrefix =
Low, Chetwin and Wang, Weimin and Katyal, Calder , year =. 2510.01284 , archivePrefix =
-
[19]
2601.03233 , archivePrefix =
HaCohen, Yoav and Brazowski, Benny and Chiprut, Nisan and Bitterman, Yaki and Kvochko, Andrew and Berkowitz, Avishai and Shalem, Daniel and Lifschitz, Daphna and Moshe, Dudu and Porat, Eitan and Richardson, Eitan and Shiran, Guy and Chachy, Itay and Chetboun, Jonathan and Fink...
-
[20]
2026 , eprint =
Native Audio-Visual Alignment for Generation , author =. 2026 , eprint =
2026
-
[21]
2026 , url =
Liu, Kai and Zheng, Yanhao and Wang, Kai and Wu, Shengqiong and Zhang, Rongjunchen and Luo, Jiebo and Hatzinakos, Dimitrios and Liu, Ziwei and Fei, Hao and Chua, Tat-Seng , booktitle =. 2026 , url =
2026
-
[22]
Huang, Ziqi and He, Yinan and Yu, Jiashuo and Zhang, Fan and Si, Chenyang and Jiang, Yuming and Zhang, Yuanhan and Wu, Tianxing and Jin, Qingyang and Chanpaisit, Nattapol and Wang, Yaohui and Chen, Xinyuan and Wang, Limin and Lin, Dahua and Qiao, Yu and Liu, Ziwei , booktitle =
-
[23]
2024 , doi =
Mao, Yuxin and Shen, Xuyang and Zhang, Jing and Qin, Zhen and Zhou, Jinxing and Xiang, Mochu and Zhong, Yiran and Dai, Yuchao , booktitle =. 2024 , doi =
2024
-
[24]
2026 , doi =
Shimada, Kazuki and Simon, Christian and Shibuya, Takashi and Takahashi, Shusuke and Mitsufuji, Yuki , booktitle =. 2026 , doi =
2026
-
[25]
Hua, Daili and Wang, Xizhi and Zeng, Bohan and Huang, Xinyi and Liang, Hao and Niu, Junbo and Chen, Xinlong and Xu, Quanqing and Zhang, Wentao , booktitle =
-
[26]
2512.21094 , archivePrefix =
Cao, Zhe and Wang, Tao and Wang, Jiaming and Wang, Yanghai and Zhang, Yuanxing and Wang, Jiahao and Chen, Jialu and Deng, Miao and Guo, Yubin and Liao, Chenxi and Zhang, Yize and Zhang, Zhaoxiang and Liu, Jiaheng , year =. 2512.21094 , archivePrefix =
-
[27]
2512.23994 , archivePrefix =
Xie, Tianxin and Lei, Wentao and Jiang, Kai and Huang, Guanjie and Zhang, Pengfei and Zhang, Chunhui and Ma, Fengji and He, Haoyu and Zhang, Han and He, Jiangshan and Wang, Jinting and Fang, Linghan and Gao, Lufei and Ablet, Orkesh and Zhang, Peihua and Hu, Ruolin and Li, Shen...
-
[28]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Benchmarking Single-Factor Physical Video-to-Audio Generation , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[29]
Oh, Hyun-Bin and Takida, Yuhta and Uesaka, Toshimitsu and Oh, Tae-Hyun and Mitsufuji, Yuki , booktitle =
-
[30]
The Journal of the Acoustical Society of America , volume =
The Timbre Toolbox: Extracting Audio Descriptors from Musical Signals , author =. The Journal of the Acoustical Society of America , volume =. 2011 , doi =
2011
-
[31]
The Journal of the Acoustical Society of America , volume =
New Method of Measuring Reverberation Time , author =. The Journal of the Acoustical Society of America , volume =. 1965 , doi =
1965
-
[32]
2000 , isbn =
Fundamentals of Acoustics , author =. 2000 , isbn =
2000
-
[33]
and Dai, Angela and Funkhouser, Thomas and Halber, Maciej and Niessner, Matthias and Savva, Manolis and Song, Shuran and Zeng, Andy and Zhang, Yinda , booktitle =
Chang, Angel X. and Dai, Angela and Funkhouser, Thomas and Halber, Maciej and Niessner, Matthias and Savva, Manolis and Song, Shuran and Zeng, Andy and Zhang, Yinda , booktitle =. 2017 , doi =
2017
-
[34]
Chen, Changan and Schissler, Carl and Garg, Sanchit and Kobernik, Philip and Clegg, Alexander and Calamia, Paul and Batra, Dhruv and Robinson, Philip and Grauman, Kristen , booktitle =
-
[35]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[36]
The Room Acoustics Absorption Coefficient Database , howpublished =. n.d. , note =
-
[37]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Towards Open-Vocabulary Audio-Visual Event Localization , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[38]
2025 , doi =
Hai, Jiarui and Wang, Helin and Guo, Weizhe and Elhilali, Mounya , booktitle =. 2025 , doi =
2025
-
[39]
2026 , howpublished =
2026
-
[40]
2025 , howpublished =
2025
-
[41]
2026 , month = apr, howpublished =
Alibaba Unveils. 2026 , month = apr, howpublished =
2026
-
[42]
2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages =
Elizalde, Benjamin and Deshmukh, Soham and Al Ismail, Mahmoud and Wang, Huaming , title =. 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages =. 2023 , publisher =
2023
-
[43]
2023 , month = nov, note =
Algorithms to Measure Audio Programme Loudness and True-Peak Audio Level , institution =. 2023 , month = nov, note =
2023
-
[44]
2023 , month = nov, note =
Loudness Normalisation and Permitted Maximum Level of Audio Signals , institution =. 2023 , month = nov, note =
2023
-
[45]
2025 , eprint =
Tjandra, Andros and Wu, Yi-Chiao and Guo, Baishan and Hoffman, John and Ellis, Brian and Vyas, Apoorv and Shi, Bowen and Chen, Sanyuan and Le, Matt and Zacharov, Nick and Wood, Carleigh and Lee, Ann and Hsu, Wei-Ning , title =. 2025 , eprint =
2025
-
[46]
and Torralba, Antonio and Adelson, Edward H
Owens, Andrew and Isola, Phillip and McDermott, Josh H. and Torralba, Antonio and Adelson, Edward H. and Freeman, William T. , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =. 2016 , doi =
2016
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.