{"id":"602bd239-8940-4081-af65-f52f98345c41","arxiv_id":"2504.13608","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Using all-to-all bidirectional consistency losses between label-tree levels, plus orthogonal enhancement of attention masks and features, gives modest accuracy and consistency gains on three fine-grained benchmarks.","lead":"A new training method helps image classifiers use the family tree of categories, like bird order, family, and species, to reduce mistakes at every level of detail. It reports small accuracy gains on three standard fine-grained classification benchmarks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CHBC's reported gains may stem from unexplained target-dataset pretraining of the predict submodule (§4.2), not from the MGE/CBC modules; no evidence shows baselines or SOTA received the same initialization.","rationale":"The reader identified this exact confound, and it is the most load-bearing issue in the paper. The central claim is that the MGE and CBC modules cause the accuracy and consistency improvements; the only evidence for that claim is empirical, so any uncontrolled advantage in the setup undermines attribution. Section 4.2 introduces a nonstandard target-dataset pretraining for the predict submodule without specifying the split, task, or duration, and without establishing that the baselines shared the same initialization. This is more directly tied to the central claim than the other listed issues (no error bars, architecture ambiguities, no code), because it could fully explain the headline margins. The rest of the method is internally coherent, the ablations are consistent with the proposed mechanism once the initialization question is controlled, and the reported gains are plausible. Therefore the correct outcome is not rejection but a conditional acceptance requiring the missing specification and a controlled rerun. Because the reader's verdict is already CONDITIONAL, my assessment leaves that verdict unchanged.","tokens_in":15365,"tokens_out":7199,"duration_ms":71164,"concrete_test":"Re-run the CUB experiments in Tables 1, 3, and 7 with the predict submodule initialized from ImageNet only, matching the attention submodule's initialization and keeping all other settings fixed. Record species accuracy, wa_acc, and TCR. If the numbers drop materially (e.g., species accuracy falls below the reported 87.8 or wa_acc below 90.4), the reported gains depend on the unexplained dataset-pretrained initialization. Additionally, train baseline-multi with the same target-dataset pretraining protocol and compare; if CHBC's margin over this controlled baseline shrinks to below 1%, the central attribution of gains to MGE and CBC fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In §4.2, the predict submodule is 'initialized with parameters pre-trained on the specific datasets in Section 4.1', attributed to HSE [16]. The paper does not state the pretraining split (train-only vs. full dataset), task, or duration, and does not report that baseline-multi, baseline-single, or the reproduced SOTA methods received the same in-domain initialization. The central claim that MGE and CBC improve accuracy and consistency rests entirely on the empirical results in Tables 1–3 and the ablations in Table 7. Because the predict submodule initialization already encodes target-dataset information, the observed gains over baseline-multi—3.1% on CUB species, 3.0% on Air models, and 2.1% on Cars models—are confounded. If the initialization provides a better starting point than the baselines' initialization, the improvements may be attributable to the initialization rather than to the proposed modules. If the pretraining uses the full dataset or test split, the results would additionally be inflated by leakage. The ablations do not resolve this because the first row of Tables 5–7 is baseline-multi, which is not described as using the same predict-submodule initialization. This is a load-bearing attribution problem, not a mere implementation detail.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CHBC, a framework for fine-grained visual classification that exploits tree-structured label hierarchies without extra annotations. The method has two main components: a Multi-Granularity Enhancement (MGE) module that decomposes and enhances attention masks and features across hierarchy levels via matrix orthogonalization, and a Cross-hierarchical Bidirectional Consistency (CBC) module that enforces consistency between predictions at different granularities using JS divergence in an all-to-all interaction scheme. Experiments are reported on CUB-200-2011, FGVC-Aircraft, and Stanford Cars, with comparisons to hierarchical multi-label methods such as HSE, FGN, HRN, and HCSL, plus ablations of the proposed components. The central claims are that CHBC improves both weighted-average accuracy and the proposed Tree-based Consistency Rate (TCR) over baselines and prior multi-level methods.","tokens_in":15623,"tokens_out":3613,"duration_ms":35601,"significance":"If the empirical results are reliable, the paper makes a modest but useful contribution: it addresses hierarchical fine-grained classification without additional part or bounding-box annotations, and it introduces a bidirectional consistency mechanism plus a TCR metric that directly measures whether predicted labels respect the hierarchy. The ablations in Tables 5-7 are internally consistent and show that each component contributes. The method also does not fit parameters to the reported metrics, and the consistency loss is derived from the externally given tree hierarchy. However, the main empirical claim depends on an undocumented dataset-pretraining step for the predict submodule and on single-run results with small margins over strong baselines, so the significance of the reported gains is currently not fully established.","major_comments":[{"comment":"The paper states that the predict submodule is initialized with parameters pre-trained on the specific datasets in Section 4.1, but it does not specify the pre-training split (training-only versus full dataset), the task, or the duration. This is load-bearing because Tables 1-3 and the ablations in Tables 5-7 compare CHBC against baseline-multi and reproduced SOTA methods, and the reader cannot tell whether those baselines received the same in-domain initialization. If the pre-training uses the test split, the reported numbers would be inflated by leakage; if it merely gives CHBC a better starting point, the gains in Tables 1-3 (e.g., 3.1% on CUB species over baseline-multi) would not be attributable to the MGE and CBC modules. The authors must report the exact pre-training protocol, apply it identically to all compared baselines, or remove this dataset-level pre-training and re-run the experiments.","section":"Section 4.2 (Implementation details)"},{"comment":"All accuracy and TCR numbers are reported from single runs with no error bars or repeated-seed statistics, while several improvements over the strongest prior method are very small: Air wa_acc is 95.3 versus 95.1 for HCSL in Table 2, CUB species accuracy is 87.8 versus 87.7 for HCSL in Table 1, and Cars maker accuracy is 97.8, which is actually below HCSL's 97.9. With such margins, the claim that CHBC consistently outperforms prior methods is not supported without variance estimates or significance testing. Please provide mean and standard deviation over at least three seeds, or otherwise demonstrate that the differences are not within run-to-run noise.","section":"Section 4.3 (Results and analysis), Tables 1-3"},{"comment":"The ablation rows labeled 'Base' are described as baseline-multi, but it is not stated whether baseline-multi uses the same predict-submodule initialization on the target datasets as CHBC. Since the full model and all ablated variants presumably share this initialization, the ablations isolate the effect of MGE and CBC only under the assumption that the initialization is constant across rows. This assumption is not documented, and if baseline-multi is initialized purely from ImageNet while CHBC is initialized from target-dataset weights, the 1.5% MGE-only and 2.3% CBC-only gains in Table 7 could reflect the initialization rather than the modules. Please clarify the initialization for every row in the ablation tables.","section":"Section 4.4 (Ablation study), Tables 5-7"}],"minor_comments":[{"comment":"The denominator in the MOD formula sums squared entries of the coarse matrix; the behavior when a coarse matrix is entirely zero at some spatial locations is not discussed. Please specify whether any numerical stabilization is used.","section":"Section 3.2, Eq. (5)"},{"comment":"The notation \"s_i × D_i,j\" is ambiguous: it should be a matrix-vector product, not entrywise multiplication. Please clarify the operation and specify the dimensions of the resulting vector.","section":"Section 3.3, Eq. (13)"},{"comment":"The sentence \"According to HSE [16], this initialization accelerates the convergence of models\" cites a prior work for the initialization practice, but HSE [16] is not cited with a page or section number, and the current paper does not describe how the pre-training is performed in practice. Please expand this description.","section":"Section 4.2"},{"comment":"The plots for hyperparameters α and T lack axis labels and numerical tick values, making it difficult to read the actual sensitivities. Please add proper axis labels and legend entries.","section":"Figure 7"},{"comment":"Reference [11] has a typo in the title: \"Fine-grained, ornot\" should be \"Fine-grained, or not\". Please also check for consistent capitalization of \"Tree Hierarchy\" throughout the text.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The central issue is the undocumented target-dataset pre-training of the predict submodule in Section 4.2; if the authors cannot show that all baselines received identical initialization, the empirical attribution of gains to MGE and CBC is not justified. The single-run results and small margins over HCSL further weaken the empirical claim. These concerns are addressable through additional experiments and clearer reporting, so I recommend major revision rather than rejection. The paper's idea is reasonable and the ablations are well structured, but the evidence in the current form is not yet convincing enough for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: CHBC is a sensible, well-ablated combination of known ingredients—orthogonal decomposition of attention masks and features, plus all-to-all bidirectional JS consistency—and the paper is worth a serious look. But the reported gains are not currently attributable to those modules, because the predict submodule starts from parameters pre-trained on the target datasets and the baselines/SOTA are not described as getting the same start.\n\nWhat's actually new: the specific MGE decomposition (using previous-level attention and first-level features) plus CBC's all-to-all bidirectional consistency loss is not in the cited hierarchical-FGVC line. The ablations in Tables 5–7 are internally consistent: MatOrth beats additive residual strategies, JS beats KL and EMD, all-to-all beats neighbor-only and finest-only, and each module helps. The TCR metric is reasonable and the comparison set is representative. No parameter is fitted to the reported metrics, and the Tree Hierarchy is external to the model. Credit where due: this is a clean architecture writeup with honest ablations.\n\nThe soft spot is the initialization, and it is load-bearing. Section 4.2 says the predict submodule is initialized \"with parameters pre-trained on the specific datasets,\" citing HSE. It doesn't say whether that pretraining used only the training split or the full dataset, what task, or for how long. More importantly, nothing says baseline-multi, baseline-single, or the reproduced HSE/FGN numbers used the same initialization. If baseline-multi starts from ImageNet weights and CHBC starts from dataset-pretrained weights, then the headline gains—3.1% on CUB species, 3.0% on Air, 2.1% on Cars—could come from the starting point, not from MGE/CBC. The ablations don't close this because the Base row is baseline-multi, which is not described as using the same predict-submodule initialization. This is not a minor implementation detail; it affects the central attribution claim.\n\nOther issues are more ordinary: no code, single runs without error bars, and several SOTA deltas of 0.1–0.3 points. Those would be tolerable with code and repeated runs; the initialization issue needs a direct answer.\n\nWho this is for: people working on hierarchical multi-granularity FGVC, especially those building on HSE, HRN, and HCSL. The architecture recipe and ablation findings are useful. It deserves a serious referee, but acceptance should be conditional on clearing up the pretraining confound and releasing code.","headline":"A clean, well-ablated combination of known hierarchical-consistency ideas, but the unexplained dataset-pretrained initialization of the predict submodule makes the reported gains unproven.","tokens_in":16131,"tokens_out":3021,"would_cite":false,"duration_ms":28139,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The central claim is that using a label tree's hierarchy as a bidirectional consistency constraint improves fine-grained classification accuracy and consistency, with reported finest-level gains of 3.1%, 3.0%, and 2.1% over a multi-label…","keywords":["fine-grained visual classification","hierarchical labels","multi-granularity","bidirectional consistency","attention masks","matrix orthogonal decomposition","Jensen-Shannon divergence","tree hierarchy"],"falsifier":"Retrain CHBC and baseline-multi with the predict branches initialized only from ImageNet weights, removing all target-dataset pretraining; if the reported 3.1% CUB species-level gain over baseline-multi shrinks to near zero, the improvement comes from the hidden in-domain initialization rather than from MGE and CBC.","tokens_in":15176,"feed_emoji":"🐦","tokens_out":9746,"duration_ms":77938,"temperature":0.7,"pith_summary":"This paper tries to show that a fine-grained classifier can get a free accuracy and consistency boost from the tree-structured label hierarchy already present in the data, without any additional annotations. The proposed CHBC framework refines each hierarchy level's attention masks and features by orthogonal decomposition, then couples every level's prediction distribution with every other level in both directions by minimizing Jensen-Shannon divergence. On three standard FGVC datasets, the reported weighted-average accuracy and Tree-based Consistency Rate are the highest among the compared multi-level methods, and finest-level accuracy rises by 3.1% (CUB), 3.0% (Aircraft), and 2.1% (Cars) over the multi-label baseline. A sympathetic reading is that hierarchical label structure, usually discarded in fine-grained recognition, is itself a useful source of supervision.","feed_headline":"Bidirectional hierarchy loss lifts fine-grained accuracy by 3.1%","feed_subtitle":"A tree-structured consistency constraint raises species-level accuracy on CUB, Aircraft, and Cars with no extra annotations.","key_machinery":"Two modules carry the argument. MGE generates a CAM attention mask and a level-specific feature map for each hierarchy level; the matrix orthogonal decomposition $M_{\\mathrm{orth}} = M_{\\mathrm{fine}} - \\frac{\\sum_{m,n} M_{\\mathrm{fine}}^{(m,n)} M_{\\mathrm{coarse}}^{(m,n)}}{\\sum_{m,n} M_{\\mathrm{coarse}}^{(m,n)} M_{\\mathrm{coarse}}^{(m,n)}} M_{\\mathrm{coarse}}$ removes the coarse component from the fine matrix and adds the residual scaled by $\\alpha$ back, applied to attention from the previous level and features from the coarsest level. CBC uses the tree adjacency matrix $D_{i,j}$ to expand coarse distributions to the fine dimension and to aggregate fine distributions to the coarse dimension, then defines the combined distribution $\\hat{s}_l$ from all other levels and minimizes the Jensen-Shannon divergence $JS(s_l, \\hat{s}_l)$ for every level $l$, including an all-level concatenated header. This all-to-all bidirectional loss is what makes the hierarchy constrain every prediction at once.","core_discovery":"The paper proposes CHBC and claims that decomposing attention masks and features across hierarchies, then enforcing bidirectional consistency among all hierarchy levels, improves both accuracy and label consistency. The key reported result is that finest-level accuracy improves by 3.1% on CUB-200-2011, 3.0% on FGVC-Aircraft, and 2.1% on Stanford Cars over the multi-label baseline, while weighted average accuracy and the proposed Tree-based Consistency Rate are the best among the compared methods on all three datasets. The authors attribute this to the Multi Granularity Enhancement (MGE) module, which uses matrix orthogonal decomposition to pull each level's discriminative content away from coarser levels, and to the Cross-hierarchical Bidirectional Consistency (CBC) module, which projects coarse predictions down to fine levels and fine predictions up to coarse levels and minimizes their JS divergence in an all-to-all interaction scheme.","pith_inferences":["The CBC loss can likely be detached from MGE and applied as a standalone regularizer to any hierarchical classifier; a clean test would add only the all-to-all JS term to baseline-multi and measure the gain.","Because CBC raises the probabilities of sibling subclasses under a confident superclass, it should systematically improve Top-3 and Top-5 accuracy, and the gains could be evaluated as a calibration improvement rather than only Top-1 accuracy.","The all-to-all scheme computes a consistency term for every level pair, so on deeper trees the cost grows; a sampled or neighbor-pruned consistency graph might recover most of the benefit at lower cost.","The reported dependence on target-dataset pretraining of the predict branches means the architecture's contribution should be re-measured under a shared initialization protocol before attributing the accuracy gains to MGE and CBC."],"forward_implications":["If the reported gains hold, hierarchical label trees improve even finest-level accuracy, so FGVC models should use existing label hierarchies instead of discarding them.","The ablations claim that all-to-all bidirectional consistency beats both neighbor-only and all-to-finest interaction, implying consistency should be enforced globally across the tree rather than only along adjacent edges.","The Tree-based Consistency Rate gains (85.0% on CUB, 92.5% on Aircraft, 94.3% on Cars) imply that when the model errs, the wrong fine label is more likely to stay inside the correct superclass, making errors less misleading in practice.","At the finest level CHBC reports the best accuracy on Aircraft and Cars among compared single-label methods and remains competitive on CUB, so using hierarchy does not appear to trade away fine-grained performance."],"supporting_citations":[{"why":"Defines the hierarchical label structure used for CUB and supplies the target-dataset pretraining of the predict branches.","marker":"[16]"},{"why":"Provides the CAM attention-mask generation that MGE builds on.","marker":"[27]"},{"why":"A hierarchical residual network baseline that CHBC is compared against and which motivates adding coarse features to the fine level.","marker":"[14]"},{"why":"A consistency-aware hierarchical method that CHBC extends and outperforms in the main comparisons.","marker":"[12]"},{"why":"Supplies the multi-granularity user-need motivation and the FGN baseline reported in the experiments.","marker":"[11]"},{"why":"CUB-200-2011 is the four-level ornithological benchmark on which the main CUB results are measured.","marker":"[28]"},{"why":"FGVC-Aircraft provides the three-level maker/family/model hierarchy used for the Air experiments.","marker":"[29]"},{"why":"Stanford Cars provides the maker/model hierarchy used for the Cars experiments.","marker":"[30]"},{"why":"ResNet-50 is the backbone split into trunk net and MGE modules.","marker":"[31]"},{"why":"ImageNet pretraining initializes the trunk net and attention submodules.","marker":"[32]"}],"fun_headline_variants":["CHBC: Bidirectional tree consistency lifts fine-grained accuracy","Tree hierarchy consistency boosts fine-grained classification by 3.1%","Cross-hierarchical bidirectional learning sharpens fine-grained visual classes","No-annotation hierarchy consistency improves fine-grained recognition","Consistency across label hierarchies reduces fine-grained misclassification"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The per-level predict branches are initialized with parameters pre-trained on the target datasets themselves, and the paper does not show that the baseline and comparison models received the same in-domain pretraining, so if that pretraining drives the gains, the proposed MGE and CBC modules may not be responsible.","fun_headline_variants_meta":{"raw":{"variants":["CHBC: Bidirectional tree consistency lifts fine-grained accuracy","Tree hierarchy consistency boosts fine-grained classification by 3.1%","Cross-hierarchical bidirectional learning sharpens fine-grained visual classes","No-annotation hierarchy consistency improves fine-grained recognition","Consistency across label hierarchies reduces fine-grained misclassification"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000222,"raw_usage":{"total_tokens":1418,"prompt_tokens":877,"completion_tokens":541,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":460}},"tokens_in":493,"tokens_out":541,"duration_ms":5110,"temperature":1.0,"reasoning_tokens":460,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:03:40.815412+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain CHBC and baseline-multi with the predict branches initialized only from ImageNet weights, removing all target-dataset pretraining; if the reported 3.1% CUB species-level gain over baseline-multi shrinks to near zero, the improvement comes from the hidden in-domain initialization rather than from MGE and CBC.","supporting_citations":[{"cited_title":"4848–4857.doi:10.1109/CVPR52688","cited_arxiv_id":null,"evidence_quote":"A hierarchical residual network baseline that CHBC is compared against and which motivates adding coarse features to the fine level."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"CUB-200-2011 is the four-level ornithological benchmark on which the main CUB results are measured."}],"review_version":1}