Pith. sign in

REVIEW 6 cited by

LLM-Prop: Predicting Physical And Electronic Properties Of Crystalline Solids From Their Text Descriptions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.14029 v1 pith:ASIQQSPA submitted 2023-10-21 cs.CL cond-mat.mtrl-sci

classification cs.CLcond-mat.mtrl-sci
keywords crystalpropertiespredictingtextdescriptionsllm-propcurrentgnns
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The prediction of crystal properties plays a crucial role in the crystal design process. Current methods for predicting crystal properties focus on modeling crystal structures using graph neural networks (GNNs). Although GNNs are powerful, accurately modeling the complex interactions between atoms and molecules within a crystal remains a challenge. Surprisingly, predicting crystal properties from crystal text descriptions is understudied, despite the rich information and expressiveness that text data offer. One of the main reasons is the lack of publicly available data for this task. In this paper, we develop and make public a benchmark dataset (called TextEdge) that contains text descriptions of crystal structures with their properties. We then propose LLM-Prop, a method that leverages the general-purpose learning capabilities of large language models (LLMs) to predict the physical and electronic properties of crystals from their text descriptions. LLM-Prop outperforms the current state-of-the-art GNN-based crystal property predictor by about 4% in predicting band gap, 3% in classifying whether the band gap is direct or indirect, and 66% in predicting unit cell volume. LLM-Prop also outperforms a finetuned MatBERT, a domain-specific pre-trained BERT model, despite having 3 times fewer parameters. Our empirical results may highlight the current inability of GNNs to capture information pertaining to space group symmetry and Wyckoff sites for accurate crystal property prediction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 29 citations worldwide. Full citation record

  1. Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Materials-science mechanisms are readable and steerable in a Gemma LLM through matched state changes, while absolute hidden-state graphs fail to uniquely encode physical polarity.

  2. AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org

    cs.AI 2025-12 reject novelty 6.0 of 10

    An open-source agentic materials-design platform shows tool access can help or hurt accuracy depending on the property, but its headline memorization-resistant test results are not presented.

  3. Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new 1,500-question benchmark shows multimodal LLMs score about 26 to 31 points below human experts on understanding materials characterization images.

  4. SiPhy: Single-Image Physical Property Reasoning

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A single-image vision-language pipeline reports state-of-the-art mass, density, and stiffness predictions by combining CLIP features, a fine-tuned VLM, and depth-adaptive pseudo-voxel sampling.

  5. TopoMAS: Large Language Model Driven Topological Materials Multiagent System

    cond-mat.mtrl-sci 2025-07 conditional novelty 4.0 of 10

    TopoMAS is a multi-agent LLM framework that automates retrieval, generation, and first-principles validation for topological materials, reporting 94.55% accuracy with a lightweight Qwen2.5-72B model.

  6. From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines

    cs.DL 2026-06 unverdicted novelty 3.0 of 10

    LLMs accelerate research workflows from idea generation to writing but introduce challenges like hallucination, bias, opacity, and ten systemic risks requiring new governance frameworks.

Pith tools