Pith. sign in

REVIEW 4 cited by

UniGenX: a unified generative foundation model that couples sequence, structure and function to accelerate scientific design across proteins, molecules and materials

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.06687 v2 pith:3Q4IHTXX submitted 2025-03-09 cs.LG cond-mat.mtrl-scics.AIphysics.bio-phphysics.chem-ph

UniGenX: a unified generative foundation model that couples sequence, structure and function to accelerate scientific design across proteins, molecules and materials

classification cs.LG cond-mat.mtrl-scics.AIphysics.bio-phphysics.chem-ph
keywords generationunigenxacrossfunctiongenerativematerialsmodelsequences
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Function in natural systems arises from one-dimensional sequences forming three-dimensional structures with specific properties. However, current generative models suffer from critical limitations: training objectives seldom target function directly, discrete sequences and continuous coordinates are optimized in isolation, and conformational ensembles are under-modeled. We present UniGenX, a unified generative foundation model that addresses these gaps by co-generating sequences and coordinates under direct functional and property objectives across proteins, molecules, and materials. UniGenX represents heterogeneous inputs as a mixed stream of symbolic and numeric tokens, where a decoder-only autoregressive transformer provides global context and a conditional diffusion head generates numeric fields steered by task-specific tokens. Besides the new high SOTAs on structure prediction tasks, the model demonstrates state-of-the-art or competitive performance for the function-aware generation across domains: in materials, it achieves "conflicted" multi-property conditional generation, yielding 436 crystal candidates meeting triple constraints, including 11 with novel compositions; in chemistry, it sets new benchmarks on five property targets and conformer ensemble generation on GEOM; and in biology, it improves success in modeling protein induced fit (RMSD < 2 {\AA}) by over 23-fold and enhances EC-conditioned enzyme design. Ablation studies and cross-domain transfer substantiate the benefits of joint discrete-continuous training, establishing UniGenX as a significant advance from prediction to controllable, function-aware generation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design

    q-bio.QM 2026-05 unverdicted novelty 7.0

    VibeProteinBench is a three-stage language-interfaced benchmark revealing that no current LLM performs strongly across recognition, engineering, and generation of proteins.

  2. VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design

    q-bio.QM 2026-05 unverdicted novelty 7.0

    VibeProteinBench is a new benchmark evaluating LLMs on open-ended language-interfaced protein design across recognition, engineering, and generation, with no model showing strong performance in all areas.

  3. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

    q-bio.QM 2026-07 conditional novelty 6.5

    A 52.6B-token multi-domain biology pretraining corpus with tool enrichment and new binding/localization instructions doubles a fixed base LLM's matched biology-eval score with little language forgetting.

  4. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

    q-bio.QM 2026-07 conditional novelty 6.0

    TheBioCollection, a 52.6B-token unified biology corpus with tool-computed text and new instruction tasks, raises a fixed 16B LLM's score on its matched biology eval from 0.223 to 0.499 (2.24×).