Pith. sign in

REVIEW 3 cited by

AdaLead: A simple and robust adaptive greedy search algorithm for sequence design

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.02141 v1 pith:Q2FSC2VU submitted 2020-10-05 cs.LG math.OCq-bio.BMq-bio.QM

classification cs.LGmath.OCq-bio.BMq-bio.QM
keywords adaleadapproachesdesignflexsmodelssequencesalgorithmalgorithms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Efficient design of biological sequences will have a great impact across many industrial and healthcare domains. However, discovering improved sequences requires solving a difficult optimization problem. Traditionally, this challenge was approached by biologists through a model-free method known as "directed evolution", the iterative process of random mutation and selection. As the ability to build models that capture the sequence-to-function map improves, such models can be used as oracles to screen sequences before running experiments. In recent years, interest in better algorithms that effectively use such oracles to outperform model-free approaches has intensified. These span from approaches based on Bayesian Optimization, to regularized generative models and adaptations of reinforcement learning. In this work, we implement an open-source Fitness Landscape EXploration Sandbox (FLEXS: github.com/samsinai/FLEXS) environment to test and evaluate these algorithms based on their optimality, consistency, and robustness. Using FLEXS, we develop an easy-to-implement, scalable, and robust evolutionary greedy algorithm (AdaLead). Despite its simplicity, we show that AdaLead is a remarkably strong benchmark that out-competes more complex state of the art approaches in a variety of biologically motivated sequence design challenges.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reinforced sequential Monte Carlo for amortised sampling

    cs.LG 2025-10 conditional novelty 6.0 of 10

    A method that trains neural samplers using SMC-collected off-policy samples and an importance-weighted replay buffer improves mode coverage on multi-modal targets.

  2. Steering Protein Language Models

    q-bio.BM 2025-07 reject novelty 6.0 of 10

    Activation steering can guide protein language models to generate and optimize sequences with higher predicted thermostability, solubility, or GFP brightness, but only in surrogate-based evaluation.

  3. Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RL

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Ctrl-DNA applies constrained reinforcement learning to a genomic language model, generating promoter and enhancer sequences whose predicted activity is high in a target cell type and suppressed in off-target cell types.

Pith tools