Pith. sign in

REVIEW 5 cited by

AstroM$^3$: A self-supervised multimodal model for astronomy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.08842 v1 pith:LJNNKLR6 submitted 2024-11-13 astro-ph.IM cs.AI

classification astro-ph.IMcs.AI
keywords modeldataclassificationclipself-supervisedaccuracyapproachastrom
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

While machine-learned models are now routinely employed to facilitate astronomical inquiry, model inputs tend to be limited to a primary data source (namely images or time series) and, in the more advanced approaches, some metadata. Yet with the growing use of wide-field, multiplexed observational resources, individual sources of interest often have a broad range of observational modes available. Here we construct an astronomical multimodal dataset and propose AstroM$^3$, a self-supervised pre-training approach that enables a model to learn from multiple modalities simultaneously. Specifically, we extend the CLIP (Contrastive Language-Image Pretraining) model to a trimodal setting, allowing the integration of time-series photometry data, spectra, and astrophysical metadata. In a fine-tuning supervised setting, our results demonstrate that CLIP pre-training improves classification performance for time-series photometry, where accuracy increases from 84.6% to 91.5%. Furthermore, CLIP boosts classification accuracy by up to 12.6% when the availability of labeled data is limited, showing the effectiveness of leveraging larger corpora of unlabeled data. In addition to fine-tuned classification, we can use the trained model in other downstream tasks that are not explicitly contemplated during the construction of the self-supervised model. In particular we show the efficacy of using the learned embeddings for misclassifications identification, similarity search, and anomaly detection. One surprising highlight is the "rediscovery" of Mira subtypes and two Rotational variable subclasses using manifold learning and dimension reduction algorithm. To our knowledge this is the first construction of an $n>2$ mode model in astronomy. Extensions to $n>3$ modes is naturally anticipated with this approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VERITAS: Towards a General-Purpose Replication Tool for Scientific Research

    cs.AI 2026-07 conditional novelty 6.0 of 10

    VERITAS, a CLI-agent replication framework, leads every reported metric on CORE-Bench Hard and ReplicationBench against matched Claude Code baselines across 65 papers in four domains.

  2. Generalization from Low- to Moderate-Resolution Spectra with Neural Networks for Stellar Parameter Estimation: A Case Study with DESI

    astro-ph.SR 2026-02 conditional novelty 6.0 of 10

    Pre-trained MLPs on LAMOST low-resolution spectra generalize to DESI medium-resolution spectra for [Fe/H] and [α/Fe], outperforming the DESI SP pipeline in zero-shot and improving with modest fine-tuning.

  3. Image-Based Multi-Survey Classification of Light Curves with a Pre-Trained Vision Transformer

    astro-ph.IM 2025-07 conditional novelty 5.0 of 10

    A shared-weights two-branch Swin Transformer that processes ZTF and ATLAS light curves jointly reaches 69.9% macro F1, outperforming single-survey models and simple fusion strategies on 21 classes.

  4. Causal Foundation Models: Disentangling Physics from Instrument Properties

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A dual-encoder contrastive model trained on star-instrument triplets learns separate stellar and instrumental latent spaces, improving few-shot prediction of stellar parameters in simulated TESS-like light curves.

  5. From stellar light to astrophysical insight: automating variable star research with machine learning

    astro-ph.IM 2025-07 unverdicted

    An invited review of machine learning for automated variable star research, covering data cleaning, variability classification, stellar parameter inference, and foundation models.

Pith tools