Pith. sign in

REVIEW 3 cited by

Towards Trustworthy and Aligned Machine Learning: A Data-centric Survey with Causality Perspectives

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.16851 v1 pith:WY7DPA7Q submitted 2023-07-31 cs.LG cs.AI

Towards Trustworthy and Aligned Machine Learning: A Data-centric Survey with Causality Perspectives

classification cs.LG cs.AI
keywords methodslearningmachinesurveytrustworthycausalityrobustnesstechniques
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The trustworthiness of machine learning has emerged as a critical topic in the field, encompassing various applications and research areas such as robustness, security, interpretability, and fairness. The last decade saw the development of numerous methods addressing these challenges. In this survey, we systematically review these advancements from a data-centric perspective, highlighting the shortcomings of traditional empirical risk minimization (ERM) training in handling challenges posed by the data. Interestingly, we observe a convergence of these methods, despite being developed independently across trustworthy machine learning subfields. Pearl's hierarchy of causality offers a unifying framework for these techniques. Accordingly, this survey presents the background of trustworthy machine learning development using a unified set of concepts, connects this language to Pearl's causal hierarchy, and finally discusses methods explicitly inspired by causality literature. We provide a unified language with mathematical vocabulary to link these methods across robustness, adversarial robustness, interpretability, and fairness, fostering a more cohesive understanding of the field. Further, we explore the trustworthiness of large pretrained models. After summarizing dominant techniques like fine-tuning, parameter-efficient fine-tuning, prompting, and reinforcement learning with human feedback, we draw connections between them and the standard ERM. This connection allows us to build upon the principled understanding of trustworthy methods, extending it to these new techniques in large pretrained models, paving the way for future methods. Existing methods under this perspective are also reviewed. Lastly, we offer a brief summary of the applications of these methods and discuss potential future aspects related to our survey. For more information, please visit http://trustai.one.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

    cs.AI 2026-05 unverdicted novelty 6.0

    Causality provides a unifying framework for resolving trade-offs in trustworthy AI by managing invariance conflicts under changes to the data-generating process.

  2. LLM Scheming Inversely Scales with Pretraining Language Coverage

    cs.AI 2026-06 reject novelty 5.0

    A Qwen3 model exhibits higher scheming scores in low-resource languages than in English and Chinese, suggesting alignment does not transfer uniformly across languages.

  3. Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

    cs.AI 2026-05 unverdicted novelty 4.0

    Causality resolves trade-offs in trustworthy AI by treating them as invariance conflicts under different data-generating process changes.