Pith. sign in

REVIEW 5 cited by

Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.07366 v1 pith:7WCTNWHI submitted 2024-03-12 cs.CV cs.LG

classification cs.CVcs.LG
keywords deyoentropyadaptationconfidencemetricplpdpredictionsbiased
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Test-time adaptation (TTA) fine-tunes pre-trained deep neural networks for unseen test data. The primary challenge of TTA is limited access to the entire test dataset during online updates, causing error accumulation. To mitigate it, TTA methods have utilized the model output's entropy as a confidence metric that aims to determine which samples have a lower likelihood of causing error. Through experimental studies, however, we observed the unreliability of entropy as a confidence metric for TTA under biased scenarios and theoretically revealed that it stems from the neglect of the influence of latent disentangled factors of data on predictions. Building upon these findings, we introduce a novel TTA method named Destroy Your Object (DeYO), which leverages a newly proposed confidence metric named Pseudo-Label Probability Difference (PLPD). PLPD quantifies the influence of the shape of an object on prediction by measuring the difference between predictions before and after applying an object-destructive transformation. DeYO consists of sample selection and sample weighting, which employ entropy and PLPD concurrently. For robust adaptation, DeYO prioritizes samples that dominantly incorporate shape information when making predictions. Our extensive experiments demonstrate the consistent superiority of DeYO over baseline methods across various scenarios, including biased and wild. Project page is publicly available at https://whitesnowdrop.github.io/DeYO/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Active Test-time Vision-Language Navigation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    ATENA uses episodic success/failure labels and a mixture entropy objective to adapt vision-language navigation policies at test time, improving REVERIE, R2R, and R2R-CE benchmarks.

  2. Can Experts Adapt Without Training? On Test-Time Modality Generalization in MVLMs

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A training-free test-time adaptation method (MoBE) routes between modality experts by entropy and adapts their prototypes/priors online, improving medical VLM accuracy by 4.3–7.2 points across benchmarks.

  3. Uncertainty-Aware Spatial Color Correlation for Low-Light Image Enhancement

    cs.CV 2025-08 conditional novelty 5.0 of 10

    U2CLLIE is a lightweight network for brightening dark images using entropy-guided dual-domain denoising and causal correlation modules, with small PSNR/SSIM gains and mixed LPIPS results.

  4. Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language Models

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    The submitted full text does not match the abstract, so the manuscript cannot be assessed as a coherent preprint.

  5. Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An online EM algorithm fits class-conditional Gaussians to the test stream from CLIP text-embedding initializations, improving test-time adaptation accuracy over prior methods on 15 benchmarks.

Pith tools