REVIEW 1 cited by
Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper studies the best practices for automatic machine learning (AutoML). While previous AutoML efforts have predominantly focused on unimodal data, the multimodal aspect remains under-explored. Our study delves into classification and regression problems involving flexible combinations of image, text, and tabular data. We curate a benchmark comprising 22 multimodal datasets from diverse real-world applications, encompassing all 4 combinations of the 3 modalities. Across this benchmark, we scrutinize design choices related to multimodal fusion strategies, multimodal data augmentation, converting tabular data into text, cross-modal alignment, and handling missing modalities. Through extensive experimentation and analysis, we distill a collection of effective strategies and consolidate them into a unified pipeline, achieving robust performance on diverse datasets.
Forward citations
Cited by 1 Pith paper
-
Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning
A user-conditioned deep learning model that combines sidewalk images with rater attributes predicts individual walkability ratings better than image-only models, and sidewalk imagery yields higher walkability scores t...
Discussion (0). Continue with ORCID to comment.