REVIEW 18 cited by
Supervised Contrastive Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Contrastive learning applied to self-supervised representation learning has seen a resurgence in recent years, leading to state of the art performance in the unsupervised training of deep image models. Modern batch contrastive approaches subsume or significantly outperform traditional contrastive losses such as triplet, max-margin and the N-pairs loss. In this work, we extend the self-supervised batch contrastive approach to the fully-supervised setting, allowing us to effectively leverage label information. Clusters of points belonging to the same class are pulled together in embedding space, while simultaneously pushing apart clusters of samples from different classes. We analyze two possible versions of the supervised contrastive (SupCon) loss, identifying the best-performing formulation of the loss. On ResNet-200, we achieve top-1 accuracy of 81.4% on the ImageNet dataset, which is 0.8% above the best number reported for this architecture. We show consistent outperformance over cross-entropy on other datasets and two ResNet variants. The loss shows benefits for robustness to natural corruptions and is more stable to hyperparameter settings such as optimizers and data augmentations. Our loss function is simple to implement, and reference TensorFlow code is released at https://t.ly/supcon.
Forward citations
Cited by 18 Pith papers
-
DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification
A multi-vehicle naturalistic benchmark finds learned driving embeddings retain driver identity under condition matching, while descriptors collapse and video re-ID is mostly route leakage.
-
A Theory of Contrastive Learning with Natural Images
For stationary image datasets and standard augmentations, the optimal contrastive representation is partial whitening of DFT power, implemented by a CNN with sinusoidal first-layer filters and a waterfilling weight al...
-
VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models
VISA improves closed-set 3D occupancy mIoU on nuScenes by using VLM instance audits as reliability-weighted semantic supervisors during training of existing world models.
-
Do Multiple Instance Learning Models Transfer?
Pretrained multiple instance learning models transfer across organs and tasks in computational pathology, and pancancer pretraining can rival slide foundation models with far less data.
-
Structure-Specific Representational Priors Causally Control the Grokking Delay
The grokking delay is causally the time to form the right feature-level representational structure, not a fixed optimization constant or a pure weight-norm effect.
-
Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control
EMG-CrossFormer, a query-based cross-attention transformer, improves sEMG-only and multimodal hand-gesture decoding on NinaPro DB2/DB3/DB7/DB10, with accelerometer fusion producing the largest gains.
-
Towards anomaly detection searches for new physics signatures including Higgs bosons with weakly supervised machine learning
HAXAD, a weakly supervised anomaly-detection search for Higgs-plus-X new physics, is extended with new embeddings and limit-setting, and on 470 fb^-1 of pseudo-data it matches or exceeds the best single cut-based limi...
-
Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
IMAGINE pre-synthesizes intent-concealed jailbreak-like instructions via iterative latent-space expansion, and DPO with that data reduces jailbreak attack success rates on Qwen2.5, Llama3.1 and Llama3.2.
-
From Bias to Behavior: Learning Bull-Bear Market Dynamics with Contrastive Modeling
A contrastive model, B4, jointly learns price and news representations split into bullish and bearish camps, claiming better trend prediction and interpretable bias dynamics.
-
Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study
A transformer encoder trained with supervised contrastive learning on Stockfish win probabilities, combined with an advantage-axis cosine score and 6-ply beam search, reaches an estimated Elo of 2593.
-
Securing Contrastive mmWave-based Human Activity Recognition against Adversarial Label Flipping
Trajectory-aware label flipping attacks degrade contrastive mmWave human activity recognition, and a selective contrastive learning defense keeps accuracy above 90% even at 40% poisoned labels.
-
SiPhy: Single-Image Physical Property Reasoning
A single-image vision-language pipeline reports state-of-the-art mass, density, and stiffness predictions by combining CLIP features, a fine-tuned VLM, and depth-adaptive pseudo-voxel sampling.
-
Robust and Generalizable Atrial Fibrillation Detection from ECG Using Time-Frequency Fusion and Supervised Contrastive Learning
An ECG AF detector combining gated time-frequency fusion with supervised contrastive learning reports accuracy gains over six baselines on AFDB/CPSC2021 and in cross-dataset transfers.
-
Paired-Sampling Contrastive Framework for Joint Physical-Digital Face Attack Detection
A paired-sampling contrastive framework unifies physical and digital face attack detection, achieving a 2.10% ACER on the 6th Face Anti-Spoofing Challenge.
-
TrajSurv: Learning Continuous Latent Trajectories from Electronic Health Records for Trustworthy Survival Prediction
TrajSurv learns continuous latent patient trajectories from irregular EHR data using an NCDE, aligns them with SOFA severity scores via time-aware contrastive learning, and uses vector-field and trajectory-clustering ...
-
Weak Supervision for Real World Graphs
WSNET integrates weak-label classification and contrastive losses to learn node representations, outperforming baselines on weakly labeled graphs.
-
Supervised Contrastive Learning for Ordinal Engagement Measurement
A supervised contrastive ordinal classifier with time-series augmentation improves minority-class recall on DAiSEE, but not overall accuracy, and the best non-contrastive baseline nearly matches it.
-
eMargin: Revisiting Contrastive Learning with Margin-Based Separation
An adaptive margin added to InfoNCE improves time series clustering metrics but hurts linear-probe classification, exposing a disconnect between clustering scores and downstream utility.
Discussion (0). Sign in to comment.