Pith. sign in

REVIEW 2 cited by

Human and AI Perceptual Differences in Image Classification Errors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.08733 v2 pith:FXPZEXLJ submitted 2023-04-18 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords differenceshumanhumansaccuracyaloneclassificationdistributionsmodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Artificial intelligence (AI) models for computer vision trained with supervised machine learning are assumed to solve classification tasks by imitating human behavior learned from training labels. Most efforts in recent vision research focus on measuring the model task performance using standardized benchmarks such as accuracy. However, limited work has sought to understand the perceptual difference between humans and machines. To fill this gap, this study first analyzes the statistical distributions of mistakes from the two sources and then explores how task difficulty level affects these distributions. We find that even when AI learns an excellent model from the training data, one that outperforms humans in overall accuracy, these AI models have significant and consistent differences from human perception. We demonstrate the importance of studying these differences with a simple human-AI teaming algorithm that outperforms humans alone, AI alone, or AI-AI teaming.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Coverage-Constrained Human-AI Cooperation with Multiple Experts

    cs.LG 2024-11 conditional novelty 6.0 of 10

    CL2DC trains a gating model to choose between AI-only, defer-to-specific-expert, and AI-plus-specific-expert actions while a penalty enforces a target fraction of AI-only decisions.

  2. Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The authors propose SatImp, a product of a saturation weight and an importance weight, and show it improves the unlearn-retain trade-off on TOFU, WMDP, and MUSE.

Pith tools