{"id":"0eb799d4-b747-4f6a-b5a5-6409e4aa3785","arxiv_id":"2504.13524","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"OBIFormer reports state-of-the-art PSNR and SSIM on oracle bone inscription denoising benchmarks while using fewer parameters than prior transformer-based methods.","lead":"The paper introduces OBIFormer, a deep-learning network that denoises images of oracle bone inscriptions by combining channel-wise attention, glyph structure extraction, and selective feature fusion. A fast and accurate denoiser for these ancient texts could help archaeologists and automated recognition systems read heavily degraded fragments.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2 evaluates PSNR/SSIM on training images (RCRN admits this; Oracle-50K has no split), so the claimed SOTA margins may reflect memorization rather than generalization.","rationale":"The reader's weakest assumption identifies the decisive issue: Table 2 reports PSNR/SSIM on training data. I reviewed the architecture and experiments with that lens. The CSAB/GSNB/SKFF design is coherent, and the efficiency comparison is informative, but the quantitative state-of-the-art claim rests on numbers that do not control for train/test overlap. RCRN is explicitly trained and tested on the same 900 pairs; Oracle-50K has no stated split for the denoising task; and the recognition experiment applies an Oracle-50K-trained denoiser to the recognition test set. The qualitative OBC306 generalization results are encouraging but unscored. These facts do not prove overfitting, but they make the claimed margins untrustworthy. A strict held-out re-run with all baselines under identical conditions would settle the question. Because the reader already conditioned acceptance on a proper held-out evaluation, my stress-test does not change the verdict.","tokens_in":16170,"tokens_out":5922,"duration_ms":53840,"concrete_test":"Re-run the Table 2 comparison under a strict held-out protocol: for Oracle-50K top-100, split by character class into disjoint train/validation/test sets (e.g., 70/10/20); for RCRN, since the original test set is unavailable, use 5-fold cross-validation over the 900 pairs. Train OBIFormer and all nine baselines from scratch on the same training folds, with identical augmentation, loss, and schedule, and report mean and standard deviation of PSNR/SSIM on the held-out folds. Also train OBIFormer for the Section 4.5 recognition experiment only on the recognition training split, never on the recognition test images. If OBIFormer's margins persist on the held-out data, the central claim is supported; if they shrink or reverse, the reported Table 2 values are an artifact of train/test overlap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 2's headline results are not generalization numbers. Section 4.1 states that the RCRN test set is not publicly available, so the authors use the training set for training, validation, and testing. Section 4.2 describes no train/test split for the Oracle-50K denoising experiments, so the same images are used for training and evaluation. With 8.35M parameters, OBIFormer can memorize input-output pairs from the 900 RCRN pairs or from the selected Oracle-50K images, and this advantage need not be shared by all baselines. The claimed margins of 1.06 dB on Oracle-50K and 0.16 dB on RCRN are therefore not established as real gains on unseen images. The recognition experiment in Section 4.5 is also confounded: OBIFormer is trained on Oracle-50K and then applied to the Oracle-50K test images used for recognition, so the denoising model may have already seen those exact images. The OBC306 results are qualitative only and cannot substitute for a held-out quantitative evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes OBIFormer, a U-shaped encoder-decoder network for denoising oracle bone inscription (OBI) images. The architecture combines channel-wise self-attention blocks (CSABs), glyph structural network blocks (GSNBs), and a selective kernel feature fusion (SKFF) module, and it is trained with a composite loss that includes PSNR and perceptual terms for both the reconstructed image and a skeleton image. The authors report state-of-the-art PSNR/SSIM results on the Oracle-50K and RCRN datasets, recognition improvements on Oracle-50K, qualitative generalization results on OBC306, and computational-efficiency comparisons. The central claim is that OBIFormer outperforms nine existing denoising methods while being faster and lighter than transformer-based alternatives. The evaluation protocol, however, uses the training set as the test set for RCRN and does not document a held-out split for Oracle-50K, which undermines the generalization claims; the recognition experiment is also potentially confounded by dataset overlap.","tokens_in":16393,"tokens_out":6713,"duration_ms":59671,"significance":"If the reported results were established on properly held-out data, OBIFormer would be a useful contribution to OBI denoising: the architecture is clearly described, the ablation studies isolate the effect of the SKFF module and the loss weights, and the efficiency comparison shows a favorable parameter/FLOP trade-off relative to CharFormer and Restormer. The paper also addresses a domain-relevant problem with scarce annotated real data. However, the significance is currently limited because the main comparative results do not demonstrate generalization, and the synthetic degradation protocol for Oracle-50K is not reproducible as described. The architectural ideas and the downstream recognition application are worth pursuing, but the experimental validation needs substantial revision before the claims can be accepted.","major_comments":[{"comment":"The RCRN evaluation is conducted by training, validating, and testing on the same 900 training pairs, as stated in Section 4.1 ('we use the training set for training, validation, and testing'). The PSNR/SSIM numbers in Table 2 therefore measure how well the methods reconstruct images they were trained on, not how well they generalize to unseen images. With 8.35M parameters and only 900 pairs, OBIFormer can memorize the training pairs, so the reported 0.16 dB and 0.019 SSIM margins over the baselines are not evidence of state-of-the-art performance. The authors should re-run the comparisons on a held-out split (or use cross-validation) and report both training and validation/test results.","section":"Section 4.1, Table 2"},{"comment":"No train/validation/test split is described for the Oracle-50K denoising experiments. Section 4.2 only mentions selecting the top-100 characters and applying STSN-based domain adaptation; it does not say how many pairs are used for training and how many for evaluation, nor whether the evaluation images are disjoint from the training images. Without a clear held-out split, the 1.06 dB and 0.063 SSIM improvements on Oracle-50K in Table 2 cannot be interpreted as generalization results. The manuscript must specify the split and the exact number of training/evaluation pairs.","section":"Section 4.2, Table 2"},{"comment":"The recognition experiment is confounded by potential overlap between the denoising training data and the recognition test data. OBIFormer is trained on the Oracle-50K dataset, and the recognition test set is drawn from the same dataset and further split 7:3 (Section 4.5). If the images used as recognition test data were also seen by the denoiser during training, the reported 3.65-5.19% accuracy gains may reflect memorization rather than improved generalization. The authors should clarify the exact subset used for denoising training and ensure that the recognition test images are excluded from denoising training, or re-run the experiment on a disjoint held-out set.","section":"Section 4.5"},{"comment":"The generalization claim on the OBC306 dataset is supported only by qualitative examples. Section 4.7 shows denoising outputs but provides no quantitative metrics (e.g., PSNR/SSIM against clean references or recognition accuracy) and no comparison with baseline methods on the same images. The Introduction's assertion that OBIFormer shows 'strong generalization ability' on OBC306 is not substantiated by the presented evidence. A quantitative evaluation on a held-out subset of OBC306 is needed to support this claim.","section":"Section 4.7, Figs. 8-9"},{"comment":"The synthetic degradation procedure for Oracle-50K is not reproducible as described. The text says only 'we utilize STSN to apply the domain adaptation,' without specifying the noise types, the degradation parameters, how the noisy input is generated from the clean handprint, or the number of synthetic pairs. Since Table 2's Oracle-50K results depend entirely on this procedure, the authors should provide a precise description of the synthesis pipeline, including the relationship to the four noise categories in Fig. 1 and the exact role of STSN.","section":"Section 4.2"}],"minor_comments":[{"comment":"Equations (12) and (13) are identical as written; the SKFF module should compute the softmax attention jointly over the reconstruction and glyph branches. Please correct the formulas to match the standard selective kernel formulation in Fig. 4(d).","section":"Section 3.2, Eqs. (12)-(13)"},{"comment":"The loss term L1 in Eq. (18) is defined as PSNR, which is higher-is-better, but Eq. (17) combines the terms with positive weights in a minimization objective. Please clarify the sign convention, e.g., by defining L1 as -PSNR or MSE.","section":"Section 3.3, Eqs. (17)-(18)"},{"comment":"The index notation in Eq. (2) is confusing: the decoder stage index n is not defined, and 'OFB_i' with 'OFB_{2N-i+1}' is inconsistent with the description. Please rewrite the equation and surrounding text more clearly.","section":"Section 3.1, Eq. (2)"},{"comment":"In the GSNB paragraph, 'the GSNB in n-th OFB' uses an undefined variable n, and 'RSAB' appears to be a typo for 'CSAB'. Please fix these issues.","section":"Section 3.2"},{"comment":"The 'Raw Image' row reports an SSIM of 0.099 on Oracle-50K, which is much lower than the SSIM of the same noisy image as processed by any method. Please verify that this value is not a typo (e.g., 0.909).","section":"Table 2"},{"comment":"The phrase 'the second two refer to the reconstructed skeleton image and its ground truth' is awkward; please use 'the next two images' for clarity.","section":"Section 4.8, Fig. 10"},{"comment":"The sentence 'For each pair, we split it into a noisy image and a clean image' is a tautology; please rephrase to 'Each pair consists of a noisy image and a clean image.'","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's core weakness is the evaluation protocol: the authors themselves disclose in Section 4.1 that RCRN is tested on the training set, and the Oracle-50K denoising experiments do not document a held-out split. This is not a presentational issue but a substantive flaw in the evidence for the state-of-the-art claim. The paper would need a substantially revised experimental setup, including disjoint training/test splits and a quantitative generalization evaluation, before it can be considered for publication. I also note that the code is only promised, not provided, and the synthetic data generation pipeline would need to be specified in detail to allow verification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Briefly: this is a sensible architecture paper that combines channel-wise self-attention (Restormer), glyph-structure blocks (CharFormer), and SKFF fusion, applied to oracle bone inscription denoising. The ablation studies give real support to each component, and the efficiency numbers (8.35M params, 20.45 GFLOPs, 10.35 ms) are honest and useful. The qualitative examples look plausible. So there is genuine value here for the document-analysis niche.\n\nThe problem is the headline result. Table 2 reports PSNR/SSIM on Oracle-50K and RCRN, but the paper says in Section 4.1 that the RCRN test set is not publicly available, so the authors use the training set for training, validation, and testing. That is not a test of generalization. For Oracle-50K, Section 4.2 does not describe any train/test split, so the numbers likely come from the same images used for training. With 8.35M parameters and only 900 RCRN pairs, memorization is a real risk, and the claimed margins (1.06 dB on Oracle-50K, 0.16 dB on RCRN) are not established as gains on unseen inputs. The recognition experiment in Section 4.5 is also confounded: the denoiser was trained on Oracle-50K and then applied to the Oracle-50K test images, so the test images were already seen during denoising training. The OBC306 generalization results are qualitative only and cannot substitute for held-out numbers.\n\nTo their credit, the authors disclose the RCRN limitation rather than hiding it, and they do not overclaim beyond their data. That is a point in their favor. The architecture itself is a recombination of known ingredients; it is not a fundamentally new idea, but the paper frames it appropriately and the ablations support the design choices. Missing error bars and the unavailability of the code (the GitHub link is a promise, not a release) further limit reproducibility.\n\nThe paper deserves a serious referee, but it needs major revision before acceptance. The minimal fix is a proper held-out evaluation: a real test set on RCRN (or a random split), a described split for Oracle-50K, error bars across runs, and code release. If those are provided, the SOTA claim can be re-evaluated. As it stands, the paper is a useful empirical study but not a reliable SOTA claim.","headline":"Solid engineering, but the evaluation protocol means the SOTA claim is not established; needs a real held-out test before it can be trusted.","tokens_in":16885,"tokens_out":1937,"would_cite":false,"duration_ms":16776,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"OBIFormer fuses glyph skeleton features with channel-wise self-attention to reach state-of-the-art oracle bone inscription denoising at a fraction of prior compute.","keywords":["oracle bone inscriptions","image denoising","channel-wise self-attention","glyph information","selective kernel feature fusion","transformer","character restoration","document image processing"],"falsifier":"A held-out test set of real OBI rubbings with clean ground-truth images would settle the claim: if OBIFormer's PSNR and SSIM gains over the next-best method shrink to within noise or reverse, the reported state-of-the-art result does not generalize beyond the training distribution.","tokens_in":1608,"feed_emoji":"🦴","tokens_out":7574,"duration_ms":88798,"temperature":0.7,"pith_summary":"The paper proposes OBIFormer, a U-shaped encoder-decoder network that cleans degraded images of oracle bone inscriptions by injecting glyph structure into a channel-wise self-attention backbone. The central claim is that this combination yields higher PSNR and SSIM than nine prior denoisers on synthetic and real OBI benchmarks, while needing fewer FLOPs and less inference time than the glyph-based transformer CharFormer. A sympathetic reader would care because cleaner character images are a practical step toward automatic recognition of the earliest Chinese script, where real rubbings are heavily degraded and expert annotation is scarce. The paper reports 16.31 dB PSNR / 0.893 SSIM on Oracle-50K and 22.19 dB / 0.969 on RCRN, beating the next best by margins of 1.06 dB and 0.16 dB PSNR respectively, and shows that denoised images improve ResNet recognition accuracy by roughly 3.7-5.2 percentage points.","feed_headline":"OBIFormer beats nine denoisers on oracle bone script","feed_subtitle":"Combines channel-wise attention and skeleton features for cleaner characters with fewer FLOPs.","key_machinery":"The central object is the OBIFormer block (OFB), a residual unit combining two channel-wise self-attention blocks (CSABs), two glyph structural network blocks (GSNBs), and a selective kernel feature fusion (SKFF) module. Channel-wise self-attention computes a transposed attention map of size $C \\times C$ instead of $HW \\times HW$, so cost scales with channels rather than spatial resolution; GSNBs are small residual CNNs that pull out skeleton-like glyph features; SKFF, borrowed from selective kernel networks, applies split-fuse-select to learn dynamic softmax weights over reconstruction and glyph features before summing them. These pieces work together to keep the denoising backbone cheap while ensuring the restored image preserves the character's stroke structure. The training loss combines PSNR and VGG perceptual losses on both the denoised image and a reconstructed skeleton image.","core_discovery":"The discovery the paper argues for is that glyph information, the skeletal stroke structure of a character, can be extracted from the noisy input itself and used as a conditioning signal for denoising, and that this can be done without the quadratic cost of vanilla spatial self-attention. OBIFormer routes features through two parallel streams: residual channel-wise self-attention blocks produce reconstruction features, while glyph structural network blocks produce skeleton features; a selective kernel feature fusion module learns per-position softmax weights for the two streams and sums them. The paper claims this design reaches the best PSNR and SSIM among all compared methods on both Oracle-50K and RCRN, restores broken strokes and removes spindle-shaped noise that other methods miss, and generalizes to the unseen OBC306 rubbing dataset. It also claims the architecture is light: 8.35M parameters, 20.45 G FLOPs, and 10.35 ms inference per $256\\times256$ image, roughly three times fewer parameters and 4.76 times faster than CharFormer.","pith_inferences":["A testable extension is to evaluate OBIFormer on a properly held-out split of RCRN or on OBC306 with ground-truth restoration labels; the paper's current protocol trains and tests on the same RCRN training set, so the true generalization margin is unknown.","Because the skeleton ground truth is produced by a morphological thinning method, the glyph stream's quality is bounded by that method; pairing OBIFormer with a learned skeletonizer could yield further gains and would separate glyph-extraction error from denoising error.","The same split-fuse-select pattern could transfer to other degraded ancient scripts, such as bronze inscriptions or Dunhuang manuscripts, where structural glyph priors are available but pixel-level noise is severe.","The paper's own conclusion notes weak performance on bone-cracked and dense-white-region noise due to few examples; a conditional diffusion model that synthesizes a balanced noise mix is an explicit direction the authors flag, and it could be benchmarked directly against OBIFormer's failure cases."],"forward_implications":["Denoised OBI images improve downstream recognition: with OBIFormer output, ResNet-18/50/152 accuracy on Oracle-50K rises 3.65, 4.42, and 5.19 percentage points over using noisy images.","The architecture's efficiency (8.35M parameters, 20.45 G FLOPs, 10.35 ms per image) makes it deployable on modest hardware for large rubbing collections.","The selective kernel fusion is load-bearing: replacing SKFF by addition or concatenation drops PSNR by 1.50 and 1.23 dB on RCRN, so the dynamic feature-selection mechanism is what the gains ride on.","Trained on only 900 RCRN pairs or on domain-adapted Oracle-50K, the model still denoises OBC306 rubbings, suggesting synthetic-to-real transfer is feasible for OBI denoising.","The reported gap over CharFormer, the closest glyph-based transformer, implies that the combination of channel-wise attention plus kernel fusion, rather than glyph information alone, accounts for the improvement."],"supporting_citations":[{"why":"DnCNN is the first deep-learning denoising baseline that OBIFormer must beat, supplying the standard residual-CNN comparison point.","marker":"[9]"},{"why":"Restormer is the efficient transformer baseline whose MDTA design OBIFormer adapts into channel-wise self-attention.","marker":"[10]"},{"why":"RCRN provides the real-world character image restoration dataset, the skeleton-extraction baseline, and one of the two main evaluation benchmarks.","marker":"[11]"},{"why":"CharFormer is the glyph-attentive transformer baseline; OBIFormer claims higher accuracy with far fewer parameters and faster inference.","marker":"[12]"},{"why":"Oracle-50K supplies the large handprint dataset used for training, evaluation, and recognition experiments.","marker":"[13]"},{"why":"STSN performs the domain adaptation that synthesizes noisy Oracle-50K images, making the handprint-to-rubbing training pipeline possible.","marker":"[14]"},{"why":"Selective kernel networks are the source of the SKFF split-fuse-select fusion strategy that the paper ablates against addition and concatenation.","marker":"[50]"},{"why":"The morphological skeleton extraction method generates the ground-truth skeleton images used to supervise the glyph branch.","marker":"[52]"}],"fun_headline_variants":["OBIFormer: fast glyph-aware denoising for oracle bones","Channel attention plus glyphs restore damaged oracle script","10ms denoising of oracle inscriptions with glyph guidance","Top PSNR and SSIM for oracle bone denoising at low cost","Glyph structure speeds oracle denoising and beats CharFormer"],"cache_read_input_tokens":19072,"weakest_assumption_plain":"The evaluation on the RCRN dataset uses the training set for testing because the public test set is unavailable, assuming that performance on the training distribution is a valid proxy for generalization to unseen real-world character images.","fun_headline_variants_meta":{"raw":{"variants":["OBIFormer: fast glyph-aware denoising for oracle bones","Channel attention plus glyphs restore damaged oracle script","10ms denoising of oracle inscriptions with glyph guidance","Top PSNR and SSIM for oracle bone denoising at low cost","Glyph structure speeds oracle denoising and beats CharFormer"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000297,"raw_usage":{"total_tokens":1720,"prompt_tokens":939,"completion_tokens":781,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":694}},"tokens_in":555,"tokens_out":781,"duration_ms":6844,"temperature":1.0,"reasoning_tokens":694,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:05:47.276147+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A held-out test set of real OBI rubbings with clean ground-truth images would settle the claim: if OBIFormer's PSNR and SSIM gains over the next-best method shrink to within noise or reverse, the reported state-of-the-art result does not generalize beyond the training distribution.","supporting_citations":[{"cited_title":"Restormer: Efficient transformer for high-resolution image restoration","cited_arxiv_id":null,"evidence_quote":"Restormer is the efficient transformer baseline whose MDTA design OBIFormer adapts into channel-wise self-attention."},{"cited_title":"Rcrn: Real-world character image restoration network via skeleton extraction","cited_arxiv_id":null,"evidence_quote":"RCRN provides the real-world character image restoration dataset, the skeleton-extraction baseline, and one of the two main evaluation benchmarks."},{"cited_title":"Charformer: A glyph fusion based attentive framework for high-precision character image denoising","cited_arxiv_id":null,"evidence_quote":"CharFormer is the glyph-attentive transformer baseline; OBIFormer claims higher accuracy with far fewer parameters and faster inference."},{"cited_title":"Self-supervised learning of orc-bert augmentator for recog- nizing few-shot oracle characters","cited_arxiv_id":null,"evidence_quote":"Oracle-50K supplies the large handprint dataset used for training, evaluation, and recognition experiments."},{"cited_title":"Unsupervised structure-texture separation network for oracle character recognition","cited_arxiv_id":null,"evidence_quote":"STSN performs the domain adaptation that synthesizes noisy Oracle-50K images, making the handprint-to-rubbing training pipeline possible."},{"cited_title":"Selectivekernel networks","cited_arxiv_id":null,"evidence_quote":"Selective kernel networks are the source of the SKFF split-fuse-select fusion strategy that the paper ablates against addition and concatenation."},{"cited_title":"Chinesecharacters strokethinningandextractionbasedonmathematicalmorphology[j]","cited_arxiv_id":null,"evidence_quote":"The morphological skeleton extraction method generates the ground-truth skeleton images used to supervise the glyph branch."}],"review_version":1}