Pith. sign in

REVIEW 1 cited by

Improved Off-policy Reinforcement Learning in Biological Sequence Design

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.04461 v2 pith:Q2ZM5GCU submitted 2024-10-06 cs.LG q-bio.BM

classification cs.LGq-bio.BM
keywords searchdeltalearningoff-policyproxysequencesbiologicalconservative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Designing biological sequences with desired properties is challenging due to vast search spaces and limited evaluation budgets. Although reinforcement learning methods use proxy models for rapid reward evaluation, insufficient training data can cause proxy misspecification on out-of-distribution inputs. To address this, we propose a novel off-policy search, $\delta$-Conservative Search, that enhances robustness by restricting policy exploration to reliable regions. Starting from high-score offline sequences, we inject noise by randomly masking tokens with probability $\delta$, then denoise them using our policy. We further adapt $\delta$ based on proxy uncertainty on each data point, aligning the level of conservativeness with model confidence. Experimental results show that our conservative search consistently enhances the off-policy training, outperforming existing machine learning methods in discovering high-score sequences across diverse tasks, including DNA, RNA, protein, and peptide design.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revisiting Non-Acyclic GFlowNets in Discrete Environments

    cs.LG 2025-02 accept novelty 7.0 of 10

    In cyclic discrete environments, GFlowNet flows are expected visit counts, and training a non-acyclic GFlowNet with the smallest expected trajectory length is equivalent to minimizing total flow.

Pith tools