Pith. sign in

REVIEW 10 cited by

DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.00902 v3 pith:5QTBSYOJ submitted 2023-10-02 cs.LG stat.ML

classification cs.LGstat.ML
keywords datainfinfluencemodelsdatacomputationalexistingfine-tuningmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Quantifying the impact of training data points is crucial for understanding the outputs of machine learning models and for improving the transparency of the AI pipeline. The influence function is a principled and popular data attribution method, but its computational cost often makes it challenging to use. This issue becomes more pronounced in the setting of large language models and text-to-image models. In this work, we propose DataInf, an efficient influence approximation method that is practical for large-scale generative AI models. Leveraging an easy-to-compute closed-form expression, DataInf outperforms existing influence computation algorithms in terms of computational and memory efficiency. Our theoretical analysis shows that DataInf is particularly well-suited for parameter-efficient fine-tuning techniques such as LoRA. Through systematic empirical evaluations, we show that DataInf accurately approximates influence scores and is orders of magnitude faster than existing methods. In applications to RoBERTa-large, Llama-2-13B-chat, and stable-diffusion-v1.5 models, DataInf effectively identifies the most influential fine-tuning examples better than other approximate influence scores. Moreover, it can help to identify which data points are mislabeled.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Cross-modal unlearning transfer in vision-language models is asymmetric, architecture-dependent, and shallow under typographic attacks; influence-guided block selection reduces the measured gap.

  2. A Reference-Free Framework for Evaluating Single-Frame ISP Pipelines

    eess.IV 2026-07 conditional novelty 6.0 of 10

    A proxy-reference network trained on synthetic camera pipelines estimates PSNR, SSIM, and LPIPS without a ground-truth reference, with LoRA fine-tuning adapting it to real pipelines.

  3. Understanding Data Influence with Differential Approximation

    cs.LG 2025-08 conditional novelty 6.0 of 10

    This paper introduces Diff-In, an influence estimator that accumulates second-order approximations of influence differences across training steps and shows strong accuracy in data cleaning, deletion, and coreset selec...

  4. Better Training Data Attribution via Better Inverse Hessian-Vector Products

    cs.LG 2025-07 conditional novelty 6.0 of 10

    ASTRA, an EKFAC-preconditioned Neumann series iteration, computes more accurate inverse Hessian-vector products and improves training data attribution scores over EKFAC baselines.

  5. Newfluence: Boosting Model interpretability and Understanding in High Dimensions

    stat.ML 2025-07 conditional novelty 6.0 of 10

    In high-dimensional regression, classical influence functions underestimate true leave-one-out influence by a per-point factor, and the proposed Newfluence estimator corrects this bias.

  6. Dataset Distillation by Influence Matching

    cs.CV 2026-07 reject novelty 5.0 of 10

    Inf-Match distills datasets by matching estimated parameter influence of real and synthetic data, reporting SOTA classification and retrieval, but with an unsupported theoretical core.

  7. Low-Perplexity LLM-Generated Sequences and Where To Find Them

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Only about 40% of low-perplexity 6-token spans generated by Pythia-6.9B can be exactly matched to The Pile, and the authors categorize matched and unmatched spans into four classes.

  8. Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization

    cs.IT 2026-03 reject novelty 4.0 of 10

    A curvature-based layer gain is converted into optimal allocation and pruning via water-filling and convex penalties, but the experiments substitute the baseline's influence scores and show only modest, mixed gains.

  9. Influence Functions for Preference Dataset Pruning

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Conjugate-gradient influence functions can mildly improve reward-model accuracy after pruning 10% of a preference dataset, but the gain is not statistically significant and gradient similarity better identifies helpfu...

  10. CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A collaborative data-selection method that scores each private sample's influence on a public anchor set and filters by a global threshold before federated learning or model merging.

Pith tools