Pro- teinzero: Self-improving protein generation via online reinforcement learning

Ziwen Wang, Jiajun Fan, Ruihan Guo, Thao Nguyen, Heng Ji, Ge Liu · 2025 · arXiv 2506.07459

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

read on arXiv browse 2 citing papers

citation-role summary

background 1 other 1

citation-polarity summary

background 1 unclear 1

representative citing papers

ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design

cs.LG · 2026-05-11 · unverdicted · novelty 6.0

ProteinOPD uses token-level on-policy distillation from multiple preference-specific teacher models into a shared student to balance competing objectives in protein design, delivering gains on targets without losing designability and an 8x speedup over RL baselines.

Pushing Biomolecular Utility-Diversity Frontiers with Supergroup Relative Policy Optimization

cs.CE · 2026-05-09 · conditional · novelty 6.0 · 2 refs

SGRPO is a GRPO-style framework that constructs set-level diversity rewards via supergroup sampling and leave-one-out redistribution to expand the utility-diversity Pareto frontier in biomolecular design tasks.

citing papers explorer

Showing 1 of 1 citing paper after filters.

ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design cs.LG · 2026-05-11 · unverdicted · none · ref 35
ProteinOPD uses token-level on-policy distillation from multiple preference-specific teacher models into a shared student to balance competing objectives in protein design, delivering gains on targets without losing designability and an 8x speedup over RL baselines.

Pro- teinzero: Self-improving protein generation via online reinforcement learning

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer