Pith. sign in

REVIEW 62 references

Ordinal regression works better when soft label targets evolve during training instead of staying fixed.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 18:25 UTC pith:G2QVHSDM

load-bearing objection Solid empirical ordinal recipe on CLIP; the dynamic-supervision story is real but only partly isolated from generic EMA self-distillation.

arxiv 2607.23575 v1 pith:G2QVHSDM submitted 2026-07-26 cs.CV cs.AI

D3O: Dynamic Distribution Distillation for Ordinal Regression

classification cs.CV cs.AI
keywords ordinal regressionlabel distribution learningself-distillationvision-language alignmentlabel enhancementcumulative distributionnoisy supervision
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Ordinal labels look clean on paper—age groups, aesthetic scores, disease stages—but they are usually cut from continuous judgments by people who disagree at the boundaries. Fixed one-hot or fixed soft targets force the model to treat those cuts as certain, which can lock in noise and imbalance. This paper argues that the supervision signal itself should move with the model: a self-distillation teacher recovers instance-level ordinal distributions from vision–language alignment, and those distributions keep refining as training proceeds. A second piece pushes cumulative rank structure into intermediate layers so shallow features share the same ordered geometry. On aesthetics, age, historical dating, and diabetic retinopathy, the approach beats strong baselines, and it degrades more slowly when labels are deliberately corrupted. The practical claim is simple: for ordered but subjective categories, adaptive soft targets beat static ones.

Core claim

D3O shows that replacing static ordinal supervision with training-driven evolution of label distributions—recovered by contrastive ordinal-aware enhancement and distilled through CDF-based cross-layer consistency—yields more accurate and noise-resistant ordinal models than methods that fit fixed labels or fixed soft targets.

What carries the argument

Dynamic distribution distillation: a contrastive ordinal-aware label enhancement (COLE) module builds an evolving teacher distribution from image–text alignment and ranking constraints, then CDF-based cross-layer distillation transfers cumulative ordinal structure from that teacher into intermediate network layers.

Load-bearing premise

The method assumes that text embeddings of ordinal class names already carry usable relative order, so aligning images to those embeddings produces trustworthy soft teachers rather than misleading priors—especially outside everyday image domains.

What would settle it

On a domain where class-name text embeddings lack ordinal structure (or under controlled adjacent-label noise), if D3O’s recovered teacher distributions stay misaligned with true ranks and accuracy/MAE no longer beat strong static-supervision baselines such as NumCLIP, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Ordinal tasks with subjective boundaries (aesthetics, medical grades, age bins) should treat soft targets as trainable objects, not fixed encodings.
  • Self-distillation can serve as label enhancement for ordered categories, not only as a regularizer on logits or features.
  • Propagating cumulative (CDF) structure across layers is a concrete way to keep intermediate features rank-consistent.
  • Gains should be largest under class imbalance and annotation noise, where static targets most strongly reinforce majority or wrong ranks.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If teacher quality depends on vision–language priors, purely visual ordinal settings without meaningful class text may need a different recovery path than COLE.
  • The same evolve-the-target idea could transfer to other discretized continuous attributes (pain scales, credit risk bins) where annotators disagree at thresholds.
  • A useful follow-up would measure how often the recovered distribution’s mode differs from the given hard label and whether those flips match human re-annotation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Circularity Check

0 steps flagged

No derivation circularity: empirical method paper with external test metrics; teacher soft-targets are training machinery, not predictions forced by construction.

full rationale

D3O proposes a training procedure (COLE-recovered teacher distributions via contrastive/rank losses, EMA self-distillation, CDF cross-layer KL) and evaluates it with held-out MAE/accuracy on four public ordinal benchmarks, plus a controlled label-noise protocol. Nothing in the claim chain equates a reported result to its inputs by definition: qt (Eqs. 6–9) is a model output used as a soft target, not a quantity algebraically identical to the test metrics; L_contra and L_rank anchor on training labels in the usual supervised way, which does not make test-set wins circular. There is no fitted scalar renamed as a prediction, no uniqueness theorem imported from overlapping authors to forbid alternatives, and no ansatz smuggled in via self-citation that forces the headline numbers. Concerns that gains may partly reflect generic EMA/soft-label robustness rather than ordinal-aware recovery are about experimental attribution and missing controls, not circular derivation. The paper is self-contained against external benchmarks; score 0.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 3 invented entities

Load-bearing content is methodological and empirical. Claims rest on standard deep-learning practice plus domain assumptions about ordinal ambiguity and VLM geometry, several hand-chosen loss coefficients, and two named modules without independent non-paper evidence.

free parameters (5)
  • Loss weights α1, α2 (LCOLE) and λ1, λ2 (LSD)
    Trade off contrastive, ranking, KL distillation, and CDF terms; values not reported but directly affect the central dynamic-supervision objective.
  • EMA momentum m for teacher update
    Controls how fast qt tracks the student; stated symbolically in §4.1 without a numeric value.
  • Contrastive temperature τ
    Scales similarities in Lcontra (§3.3); standard free hyperparameter, value unreported.
  • Supervised intermediate layer set S
    Which ViT blocks receive CDF distillation is left as a design choice without enumeration.
  • Adam learning rate and batch size = 1e-4, batch 64
    Fixed at 1e-4 and 64 (§4.1); training dynamics depend on these choices.
axioms (6)
  • domain assumption Real ordinal labels are discretizations of continuous semantics and therefore carry boundary ambiguity and annotation noise that static targets mishandle.
    Core motivation in §1 and Abstract; supported by cited DR rater-agreement study but treated as general.
  • domain assumption Text embeddings of ordinal class prompts already encode meaningful relative order in CLIP space, making contrastive alignment a valid carrier of ordinal geometry.
    Invoked in §3.3 citing OrdinalCLIP/NumCLIP-style evidence; strained on DR where zero-shot CLIP is majority-class.
  • domain assumption Cumulative distribution (threshold) representations encode ordinal consistency better than independent class probabilities for cross-layer transfer.
    Stated in §3.4 as motivation for Lcdf-dist; standard in cumulative link models but still an unproved design choice here.
  • domain assumption An EMA teacher provides a stable evolving supervision signal suitable for self-distillation without external labels.
    Framework premise in §3.2; standard self-KD practice.
  • standard math Softmax over cosine similarities in a shared image–text embedding space is a valid predictive model for ordinal classes.
    Problem definition Eqs. (1)–(2); standard CLIP-style classification head.
  • ad hoc to paper Distance-weighted hinge on recovered logits (Lrank) is sufficient to enforce unimodal ordinal structure in qt.
    Eq. (10) in §3.3; specific regularizer choice without theoretical optimality claim.
invented entities (3)
  • COLE (Contrastive Ordinal-Aware Label Enhancement) no independent evidence
    purpose: Recover instance-specific ordinal soft labels via attention over CLIP label embeddings plus ranking regularization, serving as dynamic teacher qt.
    Named module combining contrastive alignment, attention aggregation, MLP logits, and Lrank; no external validation outside this paper’s ablations.
  • CDF-based cross-layer interaction distillation no independent evidence
    purpose: Propagate cumulative ordinal structure from teacher distribution to intermediate ViT layer heads via L2 on CDFs.
    Specific distillation geometry introduced in §3.4; effectiveness shown only in-paper (Table 5).
  • Dynamic teacher distribution qt(y|x) from self-distillation no independent evidence
    purpose: Replace static one-hot/fixed soft targets with evolving supervision in Ldynamic / LSD.
    Central conceptual object of D3O; defined from model states rather than ground-truth distributions.

pith-pipeline@v1.2.0-grok45-kimik3 · 21875 in / 3649 out tokens · 59755 ms · 2026-07-30T18:25:46.419683+00:00 · methodology

0 comments
read the original abstract

Ordinal regression is widely used in scenarios where labels are discrete yet inherently ordered. In practice, however, ordinal labels are often obtained by discretizing underlying continuous semantics through subjective human judgment, resulting in ambiguous boundaries and annotation noise. Such uncertainty challenges existing methods that rely on fixed supervision targets, which may reinforce biased ordering under subjective annotations. To address this limitation, we propose D3O, a dynamic distribution distillation framework that replaces static supervision with training-driven evolution of ordinal label distributions via self-distillation. Specifically, we introduce a contrastive ordinal-aware label enhancement module that leverages vision-language alignment to recover refined label distributions capturing both inter-class ambiguity and instance-level uncertainty. Furthermore, we design a CDF-based cross-layer interaction distillation mechanism to propagate cumulative ordinal structure across network hierarchy, ensuring consistent ordinal geometry in intermediate representations. Extensive experiments on four general ordinal regression tasks demonstrate that our proposed D3O consistently outperforms existing approaches, particularly under severe class imbalance and noisy supervision. These results highlight the effectiveness of dynamic supervision in learning robust ordinal representations beyond fixed targets. The code will be publicly available.

Figures

Figures reproduced from arXiv: 2607.23575 by Chunlai Dong, Haochao Ying, Jian Wu, Yaojun Hu, Yuyang Xu.

Figure 1
Figure 1. Figure 1: Motivation of dynamic supervision for ordinal re [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 1
Figure 1. Figure 1: Under subjective uncertainty, such fixed targets can force [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the proposed D3O framework. Image and ordinal label embeddings are aligned in a shared vision–language [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Illustrating the Contrastive Ordinal Aware Label [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 5
Figure 5. Figure 5: Visualization of the recovered label distributions [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

62 extracted references · 9 canonical work pages

  1. [1]

    Shixing Chen, Caojin Zhang, Ming Dong, Jialiang Le, and Mike Rao. 2017. Using Ranking-CNN for Age Estimation. In2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 742–751. doi:10.1109/CVPR.2017.86

  2. [2]

    Yi Cheng, Haochao Ying, Renjun Hu, Jinhong Wang, Wenhao Zheng, Xiao Zhang, Danny Chen, and Jian Wu. 2023. Robust image ordinal regression with control- lable image generation. InProceedings of the Thirty-Second International Joint Conference on Artificial Intelligence(Macao, P.R.China)(IJCAI ’23). Article 70, 9 pages. doi:10.24963/ijcai.2023/70

  3. [3]

    Christopher Pal Christopher Beckham. 2017. Unimodal Probability Distributions for Deep Ordinal Classification. InProceedings of the 34th International Conference on Machine Learning (ICML)

  4. [4]

    Wei Chu and Zoubin Ghahramani. 2005. Gaussian Processes for Ordinal Regression.Journal of Machine Learning Research6, 35 (2005), 1019–1041. http://jmlr.org/papers/v6/chu05a.html

  5. [5]

    Dmitry Demidov, Abduragim Shtanchaev, Mihail Minkov Mihaylov, and Mo- hammad Almansoori. 2024. Extract More from Less: Efficient Fine-Grained Visual Recognition in Low-Data Regimes. In35th British Machine Vision Con- ference 2024, BMVC 2024, Glasgow, UK, November 25-28, 2024. BMVA. https: //papers.bmvc2024.org/0859.pdf

  6. [6]

    Zongyong Deng, Hao Liu, Yaoxing Wang, Chenyang Wang, Zekuan Yu, and Xuehong Sun. 2021. PML: Progressive Margin Loss for Long-Tailed Age Classifi- cation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10503–10512

  7. [7]

    Chunlai Dong, Haochao Ying, Qibo Qiu, Jinhong Wang, Danny Chen, and Jian Wu. 2025. Dual-level Fuzzy Learning with Patch Guidance for Image Ordinal Regression. InProceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, IJCAI-25, James Kwok (Ed.). International Joint Conferences on Artificial Intelligence Organization, 918...

  8. [8]

    Yao Du, Qiang Zhai, Weihang Dai, and Xiaomeng Li. 2024. Teach CLIP to Develop a Number Sense for Ordinal Regression. Springer-Verlag, Berlin, Heidelberg, 1–17. doi:10.1007/978-3-031-73013-9_1

  9. [9]

    Raúl Díaz and Amit Marathe. 2019. Soft Labels for Ordinal Regression. In2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4733–

  10. [10]

    Eibe Frank and Mark Hall. 2001. A Simple Approach to Ordinal Classification. InMachine Learning: ECML 2001, Luc De Raedt and Peter Flach (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 145–156

  11. [11]

    Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. 2018. Born Again Neural Networks. InProceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 1607–1616. https://proceedings.mlr.press/v80/furlanello18a.html

  12. [12]

    Bin-Bin Gao, Chao Xing, Chen-Wei Xie, Jianxin Wu, and Xin Geng. 2017. Deep Label Distribution Learning With Label Ambiguity.IEEE Transactions on Image Processing26, 6 (2017), 2825–2838. doi:10.1109/TIP.2017.2689998

  13. [13]

    Bin-Bin Gao, Hong-Yu Zhou, Jianxin Wu, and Xin Geng. 2018. Age Estimation Using Expectation of Label Distribution Learning. InProceedings of the Twenty- Seventh International Joint Conference on Artificial Intelligence (IJCAI). 712–718. doi:10.24963/ijcai.2018/99

  14. [14]

    Xin Geng. 2016. Label Distribution Learning.IEEE Transactions on Knowledge and Data Engineering28, 7 (2016), 1734–1748. doi:10.1109/TKDE.2016.2545658

  15. [15]

    Xin Geng, Chao Yin, and Zhi-Hua Zhou. 2013. Facial Age Estimation by Learning from Label Distributions.IEEE Transactions on Pattern Analysis and Machine Intelligence35, 10 (2013), 2401–2412. doi:10.1109/TPAMI.2013.51

  16. [16]

    Chi Keong Goh, Yanzhu Liu, and Adams Wai Kin Kong. 2018. A Constrained Deep Neural Network for Ordinal Regression. In2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 831–839. doi:10.1109/CVPR.2018.00093

  17. [17]

    Andrzej Grzybowski, Piotr Brona, Tomasz Krzywicki, Magdalena Gaca-Wysocka, Arleta Berlińska, and Anna Święch. 2022. Variability of Grading DR Screening Images among Non-Trained Retina Specialists.Journal of Clinical Medicine11, 11 (2022). doi:10.3390/jcm11113125

  18. [18]

    Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the Knowledge in a Neural Network. InNeurIPS Deep Learning and Representation Learning Workshop

  19. [19]

    Yutao Hu, Xiaolong Jiang, Xuhui Liu, Xiaoyan Luo, Yao Hu, Xianbin Cao, Baochang Zhang, and Jun Zhang. 2025. Hierarchical Self-Distilled Feature Learn- ing for Fine-Grained Visual Categorization.IEEE Transactions on Neural Networks and Learning Systems36, 3 (2025), 4005–4018. doi:10.1109/TNNLS.2021.3124135

  20. [20]

    Zengwei Huo, Xu Yang, Chao Xing, Ying Zhou, Peng Hou, Jiaqi Lv, and Xin Geng. 2016. Deep Age Distribution Learning for Apparent Age Estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops. 17–24

  21. [21]

    Isaksson, M

    Linda J. Isaksson, M. Pepa, M. Zaffaroni, G. Marvaso, D. Alterio, S. Volpe, G. Corrao, M. Augugliaro, A. Starzyńska, M. C. Leonardi, R. Orecchia, and B. A. Jereczek-Fossa. 2020. Machine Learning-Based Models for Prediction of Toxicity Outcomes in Radiotherapy.Frontiers in Oncology10 (June 2020), 790. doi:10.3389/ fonc.2020.00790

  22. [22]

    Shu Kong, Xiaohui Shen, Zhe Lin, Radomir Mech, and Charless Fowlkes. 2016. Photo Aesthetics Ranking Network with Attributes and Content Adaptation. In Computer Vision – ECCV 2016, Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling (Eds.). Springer International Publishing, Cham, 662–679

  23. [23]

    Jun-Tae Lee and Chang-Su Kim. 2019. Image Aesthetic Assessment Based on Pair- wise Comparison A Unified Approach to Score Regression, Binary Classification, and Personalization.2019 IEEE/CVF International Conference on Computer Vision (ICCV)(2019), 1191–1200. https://api.semanticscholar.org/CorpusID:207988557

  24. [24]

    Seon-Ho Lee and Chang-Su Kim. 2021. Deep Repulsive Clustering of Ordered Data Based on Order-Identity Decomposition. InInternational Conference on Learning Representations. https://openreview.net/forum?id=Yz-XtK5RBxB

  25. [25]

    Seon-Ho Lee, Nyeong Ho Shin, and Chang-Su Kim. 2022. Geometric Order Learn- ing for Rank Estimation. InAdvances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 27–39. https://proceedings.neurips.cc/paper_files/paper/ 2022/file/00358de35a101a372ea0412bed91...

  26. [26]

    Gil Levi and Tal Hassncer. 2015. Age and gender classification using convolu- tional neural networks. In2015 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 34–42. doi:10.1109/CVPRW.2015.7301352

  27. [27]

    Wanhua Li, Xiaoke Huang, Jiwen Lu, Jianjiang Feng, and Jie Zhou. 2021. Learning Probabilistic Ordinal Embeddings for Uncertainty-Aware Regression. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 13896–13905. arXiv:2103.13629 doi:10.1109/CVPR46437.2021.01369

  28. [28]

    Wanhua Li, Xiaoke Huang, Zheng Zhu, Yansong Tang, Xiu Li, Jie Zhou, and Jiwen Lu. 2022. OrdinalCLIP: learning rank prompts for language-guided ordinal regression. InProceedings of the 36th International Conference on Neural Informa- tion Processing Systems(New Orleans, LA, USA)(NIPS ’22). Curran Associates Inc., Red Hook, NY, USA, Article 2559, 13 pages

  29. [29]

    Wanhua Li, Jiwen Lu, Jianjiang Feng, Chunjing Xu, Jie Zhou, and Qi Tian. 2019. BridgeNet: A Continuity-Aware Probabilistic Network for Age Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  30. [30]

    Kyungsun Lim, Nyeong-Ho Shin, Young-Yoon Lee, and Chang-Su Kim. 2020. Order Learning and Its Application to Age Estimation. InInternational Conference on Learning Representations. https://openreview.net/forum?id=HygsuaNFwr

  31. [31]

    Chuang Ma, Tomoyuki Obuchi, and Toshiyuki Tanaka. 2025. Neural Collapse in Cumulative Link Models for Ordinal Regression: An Analysis with Unconstrained Feature Model. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=pjTbFuv9ET

  32. [32]

    Hongxu Ma, Han Zhou, Kai Tian, Xuefeng Zhang, Chunjie Chen, Han Li, Jihong Guan, and Shuigeng Zhou. 2026. GoR: A Unified and Extensible Generative Framework for Ordinal Regression. InThe Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=ys80cc2N5M

  33. [33]

    Zhenxing Niu, Mo Zhou, Le Wang, Xinbo Gao, and Gang Hua. 2016. Ordinal Regression with Multiple Output CNN for Age Estimation. In2016 IEEE Conference Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Dong et al. on Computer Vision and Pattern Recognition (CVPR). 4920–4928. doi:10.1109/CVPR. 2016.532

  34. [34]

    Frank Palermo, James Hays, and Alexei A. Efros. 2012. Dating Historical Color Images. InComputer Vision – ECCV 2012, Andrew Fitzgibbon, Svetlana Lazebnik, Pietro Perona, Yoichi Sato, and Cordelia Schmid (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 499–512

  35. [35]

    Dileepa Pitawela, Gustavo Carneiro, and Hsiang-Ting Chen. 2025. CLOC: Con- trastive Learning for Ordinal Classification with Multi-Margin N-pair Loss. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 15538–15548

  36. [36]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. InProceedings of the 38th Inter- national Conference on Machine Learning (Proceedings of Machi...

  37. [37]

    Vadim Ratner, Yoel Shoshan, and Tal Kachman. 2018. Learning multiple non- mutually-exclusive tasks for improved classification of inherently ordered labels. CoRRabs/1805.11837 (2018). arXiv:1805.11837 http://arxiv.org/abs/1805.11837

  38. [38]

    Ricanek and T

    K. Ricanek and T. Tesafaye. 2006. MORPH: a longitudinal image database of normal adult age-progression. In7th International Conference on Automatic Face and Gesture Recognition (FGR06). 341–345. doi:10.1109/FGR.2006.78

  39. [39]

    Rossano Schifanella, Miriam Redi, and Luca Maria Aiello. 2021. An Image Is Worth More than a Thousand Favorites: Surfacing the Hidden Beauty of Flickr Pictures.Proceedings of the International AAAI Conference on Web and Social Media9, 1 (Aug. 2021), 397–406. doi:10.1609/icwsm.v9i1.14612

  40. [40]

    Nagur Shareef Shaik, Teja Krishna Cherukuri, Adnan Masood, Ehsan Adeli, and Dong Hye Ye. 2025. Ordinal Label-Distribution Learning with Constrained Asymmetric Priors for Imbalanced Retinal Grading. InThe Second Workshop on GenAI for Health: Potential, Trust, and Policy Compliance. https://openreview. net/forum?id=hatzHLfYtF

  41. [41]

    Nyeong-Ho Shin, Seon-Ho Lee, and Chang-Su Kim. 2022. Moving Window Re- gression: A Novel Approach to Ordinal Regression. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18760–18769

  42. [42]

    Zichang Tan, Jun Wan, Zhen Lei, Ruicong Zhi, Guodong Guo, and Stan Z. Li

  43. [43]

    Chen, and Jian Wu

    Jinhong Wang, Jintai Chen, Jian Liu, Dongqi Tang, Danny Z. Chen, and Jian Wu

  44. [44]

    Jinhong Wang, Yi Cheng, Jintai Chen, TingTing Chen, Danny Chen, and Jian Wu

  45. [45]

    Rui Wang, Peipei Li, Huaibo Huang, Chunshui Cao, Ran He, and Zhaofeng He

  46. [46]

    Changsong Wen, Xin Zhang, Xingxu Yao, and Jufeng Yang. 2023. Ordinal Label Distribution Learning. In2023 IEEE/CVF International Conference on Computer Vision (ICCV). 23424–23434. doi:10.1109/ICCV51070.2023.02146

  47. [47]

    Xin Wen, Biying Li, Haiyun Guo, Zhiwei Liu, Guosheng Hu, Ming Tang, and Jinqiao Wang. 2020. Adaptive Variance Based Label Distribution Learning for Facial Age Estimation. InComputer Vision – ECCV 2020: 16th European Confer- ence, Proceedings, Part XXIII(Glasgow, United Kingdom). Springer-Verlag, Berlin, Heidelberg, 379–395. doi:10.1007/978-3-030-58592-1_23

  48. [48]

    Di Wu, Pengfei Chen, Xuehui Yu, Guorong Li, Zhenjun Han, and Jianbin Jiao

  49. [49]

    Ning Xu, Yun-Peng Liu, and Xin Geng. 2021. Label Enhancement for Label Distribution Learning.IEEE Transactions on Knowledge and Data Engineering33, 4 (2021), 1632–1643. doi:10.1109/TKDE.2019.2947040

  50. [50]

    InProceedings of the 37th International Conference on Neural Information Processing Systems(New Orleans, LA, USA)(NIPS ’23)

    Learning-to-rank meets language: boosting language-driven ordering align- ment for ordinal classification. InProceedings of the 37th International Conference on Neural Information Processing Systems(New Orleans, LA, USA)(NIPS ’23). Curran Associates Inc., Red Hook, NY, USA, Article 3361, 15 pages

  51. [51]

    Chuanguang Yang, Zhulin An, Helong Zhou, Linhang Cai, Xiang Zhi, Jiwen Wu, Yongjun Xu, and Qian Zhang. 2022. MixSKD: Self-Knowledge Distillation from Mixup for Image Recognition. InComputer Vision – ECCV 2022, Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (Eds.). Springer Nature Switzerland, Cham, 534–551

  52. [52]

    Xu Yang, Bin-Bin Gao, Chao Xing, Zeng-Wei Huo, Xiu-Shen Wei, Ying Zhou, Jianxin Wu, and Xin Geng. 2015. Deep Label Distribution Learning for Apparent Age Estimation. InProceedings of the IEEE International Conference on Computer Vision (ICCV) Workshops

  53. [53]

    Yuzhe Yang, Liwu Xu, Leida Li, Nan Qie, Yaqian Li, Peng Zhang, and Yandong Guo. 2022. Personalized Image Aesthetics Assessment With Rich Attributes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 19861–19869

  54. [54]

    InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)

    Spatial Self-Distillation for Object Detection with Inaccurate Bounding Boxes. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 6855–6865

  55. [55]

    Qinghai Zheng, Jihua Zhu, and Haoyu Tang. 2023. Label Information Bottleneck for Label Enhancement. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 7497–7506. doi:10.1109/CVPR52729.2023. 00724

  56. [56]

    Ning Xu, Jun Shu, Yun-Peng Liu, and Xin Geng. 2020. Variational Label Enhance- ment. InProceedings of the 37th International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 119). 10597–10606

  57. [57]

    Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022. Learning to Prompt for Vision-Language Models.Int. J. Comput. Vision130, 9 (Sept. 2022), 2337–2348. doi:10.1007/s11263-022-01653-1

  58. [60]

    Sukmin Yun, Jongjin Park, Kimin Lee, and Jinwoo Shin. 2020. Regularizing Class- Wise Predictions via Self-Knowledge Distillation. In2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 13873–13882. doi:10.1109/ CVPR42600.2020.01389

  59. [62]

    Appling, Heather E

    Wei Zhi, Alison P. Appling, Heather E. Golden, Joel Podgorski, and Li Li. 2024. Deep learning for water quality.Nature Water2, 3 (March 2024), 228–241. doi:10. 1038/s44221-024-00202-z

  60. [2018]

    doi:10.1109/TPAMI.2017.2779808

    Efficient Group-n Encoding and Decoding for Facial Age Estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence40, 11 (2018), 2610–2623. doi:10.1109/TPAMI.2017.2779808

  61. [2025]

    arXiv:2503.00952 [cs.CV] https://arxiv.org/abs/2503.00952

    A Survey on Ordinal Regression: Applications, Advances and Prospects. arXiv:2503.00952 [cs.CV] https://arxiv.org/abs/2503.00952

  62. [4742]

    doi:10.1109/CVPR.2019.00487