REVIEW 22 cited by
Data Filtering Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large training sets have become a cornerstone of machine learning and are the foundation for recent advances in language modeling and multimodal learning. While data curation for pre-training is often still ad-hoc, one common paradigm is to first collect a massive pool of data from the Web and then filter this candidate pool down to an actual training set via various heuristics. In this work, we study the problem of learning a data filtering network (DFN) for this second step of filtering a large uncurated dataset. Our key finding is that the quality of a network for filtering is distinct from its performance on downstream tasks: for instance, a model that performs well on ImageNet can yield worse training sets than a model with low ImageNet accuracy that is trained on a small amount of high-quality data. Based on our insights, we construct new data filtering networks that induce state-of-the-art image-text datasets. Specifically, our best performing dataset DFN-5B enables us to train state-of-the-art CLIP models for their compute budgets: among other improvements on a variety of tasks, a ViT-H trained on our dataset achieves 84.4% zero-shot transfer accuracy on ImageNet, out-performing models trained on other datasets such as LAION-2B, DataComp-1B, or OpenAI's WIT. In order to facilitate further research in dataset design, we also release a new 2 billion example dataset DFN-2B and show that high performance data filtering networks can be trained from scratch using only publicly available data.
Forward citations
Cited by 22 Pith papers
-
Revisiting Model Stitching In the Foundation Model Era
Final-feature matching makes heterogeneous VFMs reliably stitchable; stitched models fuse complementary strengths and enable efficient multi-VFM Stitch Trees for multimodal LLMs.
-
Meta CLIP 2: A Worldwide Scaling Recipe
A data curation and training recipe that scales CLIP from English-only data to 300+ languages from scratch, breaking the curse of multilinguality at ViT-H/14 scale.
-
Synthetic Visual Genome
A GPT-4V/GPT-4o pipeline for completing and refining scene graph annotations yields a dense synthetic dataset that, after instruction tuning, gives a 3B model strong relationship understanding and grounding results.
-
Twins: Learn to Predict Unified Representations with Focal Loss
Channel-wise concatenation of SigLIP2 and Flux VAE features into one token, trained with a focal-style flow-matching loss, yields a unified representation with 1.59 gFID on ImageNet 256 and VAE-level reconstruction.
-
Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
Object-aware soft merging of post-projector visual tokens preserves MaxSim-selectable evidence, yielding >93% token reduction and higher R@1 than full ColPali on Flickr30K and MSCOCO.
-
Distribution Matching Distillation Meets Reinforcement Learning
Combining DMD distillation with RL during training produces few-step text-to-image models that outperform their multi-step teacher on several benchmarks.
-
EgoPrivacy: What Your First-Person Camera Says About You?
A new benchmark and attack show that egocentric videos leak wearer demographics and identity well above chance, and that matching third-person clips strengthens demographic attacks.
-
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
Fine-tuning CLIP's visual encoder to match DINOv2's kernel-based similarity structure improves its fine-grained visual perception while preserving its alignment to text.
-
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
SurgVLM, a family of surgical vision-language models trained on 1.81M frames and 7.79M conversations, outperforms 14 commercial VLMs on a six-dataset surgical benchmark.
-
SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning
SCIZOR filters suboptimal and redundant state-action pairs from robot demonstrations without human labels, improving imitation-learning policy success rates by about 15% on average.
-
TokBench: Evaluating Your Visual Tokenizer before Visual Generation
TokBench measures text recognition accuracy and face similarity on reconstructed images and videos across 16 tokenizers, showing small text and faces are poorly preserved and often missed by traditional metrics.
-
Qwen-Audio-VAE Technical Report
A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.
-
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
MagicVL-2B is a 2B vision-language model for mobile phones that claims state-of-the-art-matching accuracy at 41.1% lower on-device power, via a lightweight encoder, dynamic resolution, and curriculum learning.
-
Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
Running multiple short annealing runs at different token scales can reveal per-source utility scaling curves that change data-source rankings compared with single point estimates.
-
(Almost) Free Modality Stitching of Foundation Models
A hypernetwork that generates connector weights for all image-text model pairs can rank pairs like grid search at about 10x lower training cost, but the best connector lags grid search by a few points.
-
Image Reconstruction as a Tool for Feature Analysis
Reconstruction fidelity from vision encoder features reveals which models preserve more visual information, and linear feature-space operations can produce color edits.
-
DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing
Dynamic Semantic Channel Hashing (DSCH) replaces fixed-width SCH channels with continuous, similarity-dependent widths and positions, improving tie-aware mAP on most cross- and intra-modal retrieval tasks.
-
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
OpenVision 2 shows that a caption-only generative objective can match contrastive learning for multimodal vision encoders at lower training cost, scaling to 1B parameters.
-
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
A three-stage coarse-to-fine training recipe for vision backbones produces consistent benchmark gains for lightweight multimodal LLMs.
-
AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models
AMMKD claims large gains from adaptively weighted two-teacher CLIP distillation, but its equations are internally inconsistent, its baselines are unverifiable, and its tests do not match its stated retrieval goal.
-
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
The submission cannot be reviewed as a coherent paper: its abstract and full text are two different papers, so the abstract's claims have no supporting body.
-
Generalizing vision-language models to novel domains: A comprehensive survey
A survey of VLM generalization literature organized by transferred module, with benchmark tables and a review of multimodal LLMs.
Discussion (0). Sign in to comment.