Pith. sign in

REVIEW 9 cited by

Data-centric Artificial Intelligence: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.10158 v3 pith:HRYFARTC submitted 2023-03-17 cs.LG cs.AIcs.DB

classification cs.LGcs.AIcs.DB
keywords datadata-centricsurveyartificialbuildingdevelopmentdiscussintelligence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Artificial Intelligence (AI) is making a profound impact in almost every domain. A vital enabler of its great success is the availability of abundant and high-quality data for building machine learning models. Recently, the role of data in AI has been significantly magnified, giving rise to the emerging concept of data-centric AI. The attention of researchers and practitioners has gradually shifted from advancing model design to enhancing the quality and quantity of the data. In this survey, we discuss the necessity of data-centric AI, followed by a holistic view of three general data-centric goals (training data development, inference data development, and data maintenance) and the representative methods. We also organize the existing literature from automation and collaboration perspectives, discuss the challenges, and tabulate the benchmarks for various tasks. We believe this is the first comprehensive survey that provides a global view of a spectrum of tasks across various stages of the data lifecycle. We hope it can help the readers efficiently grasp a broad picture of this field, and equip them with the techniques and further research ideas to systematically engineer data for building AI systems. A companion list of data-centric AI resources will be regularly updated on https://github.com/daochenzha/data-centric-AI

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Component Testing: Validating Agentic AI Systems

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Agentic AI cannot be adequately validated by component tests alone; trajectory-in-context validation is required, and current practice is mature only for behavioral evaluation.

  2. DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    DataClaw0 introduces an agentic data-tailoring paradigm, a 9B model trained on a synthetically generated dataset, and a new benchmark, claiming improved downstream adaptation in video generation, VQA, and GUI navigati...

  3. "Skill Issues'': Data-Centric Optimization of Lakehouse Agents

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    Data-centric optimization of skills for agents on a branching lakehouse improves accuracy by 31.9% on 25 tasks via state-verification evaluation.

  4. Feature Shift Localization Network

    cs.LG 2025-06 conditional novelty 6.0 of 10

    FSL-Net localizes shifted features between two datasets using a network trained on 1,350 datasets, matching DataFix's F1 while being about 36x faster on average.

  5. ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization

    cs.HC 2025-05 conditional novelty 6.0 of 10

    A drag-and-link interface lets users define rules for actions from body-part and object relations, generating frame labels to train temporal action localization models.

  6. Laplace Sample Information: Data Informativeness Through a Bayesian Lens

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LSI ranks training samples by informativeness using the KL divergence between Laplace-approximated posteriors with and without each sample, and the ordering transfers from a small probe to larger models.

  7. DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

    cs.SE 2026-07 conditional novelty 5.0 of 10

    DataFlow-Harness builds editable data pipelines as validated DAGs via an LLM agent, hitting 93.3% task pass rate with 72.5% lower cost than vanilla script generation.

  8. MEGG: Replay via Maximally Extreme GGscore in Incremental Learning for Neural Recommendation Models

    cs.IR 2025-09 conditional novelty 5.0 of 10

    A gradient-alignment influence score (GGscore) that selects the highest- and lowest-scoring old interactions for replay improves incremental neural recommendation slightly over random replay, mainly at large replay ratios.

  9. Importance of User Control in Data-Centric Steering for Healthcare Experts

    cs.HC 2025-05 conditional novelty 5.0 of 10

    Healthcare experts who manually adjusted training data improved a diabetes prediction model more than those using automated corrections, without losing trust or understanding.

Pith tools