Pith. sign in

REVIEW 18 cited by

Supervised Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.11362 v5 pith:EN74J2XY submitted 2020-04-23 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords contrastivelosslearningbatchclustersself-supervisedsupconsupervised
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastive learning applied to self-supervised representation learning has seen a resurgence in recent years, leading to state of the art performance in the unsupervised training of deep image models. Modern batch contrastive approaches subsume or significantly outperform traditional contrastive losses such as triplet, max-margin and the N-pairs loss. In this work, we extend the self-supervised batch contrastive approach to the fully-supervised setting, allowing us to effectively leverage label information. Clusters of points belonging to the same class are pulled together in embedding space, while simultaneously pushing apart clusters of samples from different classes. We analyze two possible versions of the supervised contrastive (SupCon) loss, identifying the best-performing formulation of the loss. On ResNet-200, we achieve top-1 accuracy of 81.4% on the ImageNet dataset, which is 0.8% above the best number reported for this architecture. We show consistent outperformance over cross-entropy on other datasets and two ResNet variants. The loss shows benefits for robustness to natural corruptions and is more stable to hyperparameter settings such as optimizers and data augmentations. Our loss function is simple to implement, and reference TensorFlow code is released at https://t.ly/supcon.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 117 citations worldwide. Full citation record

  1. DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

    cs.LG 2026-07 conditional novelty 7.0 of 10

    A multi-vehicle naturalistic benchmark finds learned driving embeddings retain driver identity under condition matching, while descriptors collapse and video re-ID is mostly route leakage.

  2. A Theory of Contrastive Learning with Natural Images

    cs.CV 2026-07 conditional novelty 7.0 of 10

    For stationary image datasets and standard augmentations, the optimal contrastive representation is partial whitening of DFT power, implemented by a CNN with sinusoidal first-layer filters and a waterfilling weight al...

  3. VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    VISA improves closed-set 3D occupancy mIoU on nuScenes by using VLM instance audits as reliability-weighted semantic supervisors during training of existing world models.

  4. Do Multiple Instance Learning Models Transfer?

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Pretrained multiple instance learning models transfer across organs and tasks in computational pathology, and pancancer pretraining can rival slide foundation models with far less data.

  5. Structure-Specific Representational Priors Causally Control the Grokking Delay

    cs.LG 2026-07 conditional novelty 6.5 of 10

    The grokking delay is causally the time to form the right feature-level representational structure, not a fixed optimization constant or a pure weight-norm effect.

  6. Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control

    cs.LG 2026-07 conditional novelty 6.0 of 10

    EMG-CrossFormer, a query-based cross-attention transformer, improves sEMG-only and multimodal hand-gesture decoding on NinaPro DB2/DB3/DB7/DB10, with accelerometer fusion producing the largest gains.

  7. Towards anomaly detection searches for new physics signatures including Higgs bosons with weakly supervised machine learning

    hep-ex 2026-07 conditional novelty 6.0 of 10

    HAXAD, a weakly supervised anomaly-detection search for Higgs-plus-X new physics, is extended with new embeddings and limit-setting, and on 470 fb^-1 of pseudo-data it matches or exceeds the best single cut-based limi...

  8. Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks

    cs.CL 2025-08 conditional novelty 6.0 of 10

    IMAGINE pre-synthesizes intent-concealed jailbreak-like instructions via iterative latent-space expansion, and DPO with that data reduces jailbreak attack success rates on Qwen2.5, Llama3.1 and Llama3.2.

  9. From Bias to Behavior: Learning Bull-Bear Market Dynamics with Contrastive Modeling

    cs.LG 2025-07 reject novelty 6.0 of 10

    A contrastive model, B4, jointly learns price and news representations split into bullish and bearish camps, claiming better trend prediction and interpretable bias dynamics.

  10. Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A transformer encoder trained with supervised contrastive learning on Stockfish win probabilities, combined with an advantage-axis cosine score and 6-ply beam search, reaches an estimated Elo of 2593.

  11. Securing Contrastive mmWave-based Human Activity Recognition against Adversarial Label Flipping

    cs.CR 2026-08 conditional novelty 5.0 of 10

    Trajectory-aware label flipping attacks degrade contrastive mmWave human activity recognition, and a selective contrastive learning defense keeps accuracy above 90% even at 40% poisoned labels.

  12. SiPhy: Single-Image Physical Property Reasoning

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A single-image vision-language pipeline reports state-of-the-art mass, density, and stiffness predictions by combining CLIP features, a fine-tuned VLM, and depth-adaptive pseudo-voxel sampling.

  13. Robust and Generalizable Atrial Fibrillation Detection from ECG Using Time-Frequency Fusion and Supervised Contrastive Learning

    q-bio.QM 2026-01 conditional novelty 5.0 of 10

    An ECG AF detector combining gated time-frequency fusion with supervised contrastive learning reports accuracy gains over six baselines on AFDB/CPSC2021 and in cross-dataset transfers.

  14. Paired-Sampling Contrastive Framework for Joint Physical-Digital Face Attack Detection

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A paired-sampling contrastive framework unifies physical and digital face attack detection, achieving a 2.10% ACER on the 6th Face Anti-Spoofing Challenge.

  15. TrajSurv: Learning Continuous Latent Trajectories from Electronic Health Records for Trustworthy Survival Prediction

    cs.LG 2025-08 conditional novelty 5.0 of 10

    TrajSurv learns continuous latent patient trajectories from irregular EHR data using an NCDE, aligns them with SOFA severity scores via time-aware contrastive learning, and uses vector-field and trajectory-clustering ...

  16. Weak Supervision for Real World Graphs

    cs.LG 2025-06 conditional novelty 5.0 of 10

    WSNET integrates weak-label classification and contrastive losses to learn node representations, outperforming baselines on weakly labeled graphs.

  17. Supervised Contrastive Learning for Ordinal Engagement Measurement

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A supervised contrastive ordinal classifier with time-series augmentation improves minority-class recall on DAiSEE, but not overall accuracy, and the best non-contrastive baseline nearly matches it.

  18. eMargin: Revisiting Contrastive Learning with Margin-Based Separation

    cs.LG 2025-07 reject novelty 4.0 of 10

    An adaptive margin added to InfoNCE improves time series clustering metrics but hurts linear-probe classification, exposing a disconnect between clustering scores and downstream utility.

Pith tools