REVIEW 4 cited by
AstroLLaVA: towards the unification of astronomical data and natural language
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We present AstroLLaVA, a vision language model for astronomy that enables interaction with astronomical imagery through natural dialogue. By fine-tuning the LLaVA model on a diverse dataset of $\sim$30k images with captions and question-answer pairs sourced from NASA's `Astronomy Picture of the Day', the European Southern Observatory, and the NASA/ESA Hubble Space Telescope, we create a model capable of answering open-ended questions about astronomical concepts depicted visually. Our two-stage fine-tuning process adapts the model to both image captioning and visual question answering in the astronomy domain. We demonstrate AstroLLaVA's performance on an astronomical visual question answering benchmark and release the model weights, code, and training set to encourage further open source work in this space. Finally, we suggest a roadmap towards general astronomical data alignment with pre-trained language models, and provide an open space for collaboration towards this end for interested researchers.
Forward citations
Cited by 4 Pith papers
-
Generalist Vision-Language Models for Fast Radio Burst detection: a zero-shot benchmark against a specialized detector
Zero-shot Gemma 4 VLMs achieve 93.65% accuracy on simulated FRB detection, statistically indistinguishable from the specialized SwinYNet detector at the default threshold, while rejecting structured RFI significantly better.
-
Semantic search for 100M+ galaxy images using AI-generated captions
A CLIP-style model trained on VLM-written galaxy captions retrieves spirals, mergers, and lenses from text queries far better than image-similarity search, with VLM re-ranking roughly doubling rare-lens recall in the top 100.
-
Exploring Image-Text Alignment for Radio Galaxy Morphologies
Text captions of radio galaxy images can classify FR-I vs FR-II morphologies comparably to image embeddings, but LoRA fine-tuning improves local class coherence without improving global image-text alignment.
-
Applying multimodal learning to Classify transient Detections Early (AppleCiDEr) I: Data set, methods, and infrastructure
AppleCiDEr combines photometry, images, metadata, and spectra in one deep learning pipeline to classify ZTF transients and variable stars, with high accuracy on common classes but poor performance on tidal disruption events.
Discussion (0). Continue with ORCID to comment.