Pith. sign in

REVIEW 7 cited by

Debiasing Vision-Language Models via Biased Prompts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.00070 v2 pith:2M2CIAKO submitted 2023-01-31 cs.LG cs.CV

classification cs.LGcs.CV
keywords modelsvision-languagedebiasinggenerativeapproachbiasedbiasesclassifiers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning models have been shown to inherit biases from their training datasets. This can be particularly problematic for vision-language foundation models trained on uncurated datasets scraped from the internet. The biases can be amplified and propagated to downstream applications like zero-shot classifiers and text-to-image generative models. In this study, we propose a general approach for debiasing vision-language foundation models by projecting out biased directions in the text embedding. In particular, we show that debiasing only the text embedding with a calibrated projection matrix suffices to yield robust classifiers and fair generative models. The proposed closed-form solution enables easy integration into large-scale pipelines, and empirical results demonstrate that our approach effectively reduces social bias and spurious correlation in both discriminative and generative vision-language models without the need for additional data or training.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding Projection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PRISM debiases CLIP by using an LLM to generate biased scene descriptions and then learning a linear projection of the embedding space that reduces spurious correlations, yielding higher worst-group accuracy on Waterb...

  2. Multi-Group Proportional Representation for Text-to-Image Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    The authors apply the MPR metric (an integral probability metric) to text-to-image generation, derive tractable forms for linear and decision-tree function classes, and use it as a fine-tuning objective that reduces i...

  3. Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A per-category best-of-K RL reward, multi-axis max@K, shifts SD3.5-M perceived-appearance distributions toward uniform coverage (Fairness Score +0.23 to +0.36) without quality loss.

  4. FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    Zero-shot vision-language models are unreliable and vary widely for depression screening, and explainability-based fairness interventions often trade away accuracy without reliable fairness gains.

  5. BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models

    cs.AI 2025-11 conditional novelty 5.0 of 10

    BioPro uses orthogonal projection on a gender-variation subspace to selectively debias vision-language models, reducing gender bias in neutral contexts while preserving explicit gender cues.

  6. Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    DualFairVL jointly debiases CLIP's text and image branches via text-guided prompts, cross-attention, a hypernetwork, and prototype losses, reporting state-of-the-art AUC and fairness (DEOdds, DPD) on eight medical ima...

  7. Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models

    cs.AI 2025-06 reject novelty 3.0 of 10

    A 4B-parameter model is claimed to explain its own reasoning through inverse attention analysis, but the paper offers no consistent evidence or artifacts.

Pith tools