REVIEW 5 major objections 8 minor 1 cited by
De Novo Generation of Hit-like Molecules from Gene Expression Profiles via Deep Learning
T0 review · 5 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A VAE-LSTM turns gene expression profiles into hit-like molecules
desk verdict A plausible VAE-LSTM approach to gene-expression-conditioned molecule generation, but the evaluation overclaims because of max-Tanimoto selection and missing conditioning controls. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the compressed condition vector $F_{G_x} = \mathrm{Encoder}(G)$, where $G$ is a gene expression profile over 978 landmark genes and the latent dimension is 64. A VAE is trained first with a $\beta$-weighted evidence-lower-bound loss, reconstruction plus KL divergence, so that the latent space captures chemically induced transcriptomic responses; only the encoder is kept for generation. The generator is an LSTM that, at each time step, concatenates the embedding of the current SMILES token with $F_{G_x}$ and predicts the next token by negative log likelihood. This concatenation makes the model conditional: the same latent vector is present throughout the autoregressive sequence, so the transcriptomic state steers which SMILES string is produced. This contrasts with the TRIOMPHE baseline, which uses transcriptomic correlation only to choose a source molecule and then generates without the profile as a condition.
What would settle it
Run HVL2Mol on a target protein's profile, then rerun the same LSTM with a latent vector drawn from an unrelated protein or from random noise; if the mean Tanimoto similarity to the target's known ligands does not fall clearly below the target-condition value over 1,000 generated molecules per condition, the gene-expression condition is not causing the hit-likeness.
Extended reading notes
Core claim
The paper's central claim is that a gene expression profile, compressed by a VAE into a 64-dimensional latent vector, is a sufficient conditioning signal for an LSTM to generate SMILES strings whose products are hit-like for the profiled condition. Trained on 13,755 chemically induced MCF7 profiles from LINCS, HVL2Mol generates molecules with 88.6% validity, 83.0% uniqueness, and 99.7% novelty. For eight knockdown and two overexpression target-protein profiles, the best Tanimoto similarities to known ligands exceed both baselines for six of the eight knockdown targets and for both overexpression targets, with MTOR and PIK3CA second only to TRIOMPHE. For disease reversal profiles from CREEDS, the model reports Tanimoto similarities of 0.58, 0.60, and 0.53 to approved drugs for gastric cancer, Alzheimer's disease, and atopic dermatitis, above the DRAGONET baseline. The claim is that transcriptomic response alone can guide de novo generation toward molecules with potential bioactivities, without known ligand structures or target 3D structures.
Load-bearing premise
The load-bearing premise is that the compressed representation learned from 978-gene chemically induced profiles in one breast-cancer cell line also represents target-protein knockdown and overexpression profiles and patient disease profiles, even though those inputs come from different gene sets and biological experiments.
Editorial extensions
If this is right
- Target-protein perturbation profiles become a sufficient starting point: with a knockdown or overexpression profile, HVL2Mol outputs candidate SMILES strings without any known ligand or receptor structure.
- Generated molecules are mostly novel (99.7%) yet retain drug-like and synthesizable properties, with average QED 0.61 versus 0.60 in training, so the generator explores near known chemical space rather than copying training compounds.
- Disease reversal profiles can be translated into therapeutic candidates: the case studies report Tanimoto similarities of 0.58, 0.60, and 0.53 to approved drugs for gastric cancer, Alzheimer's disease, and atopic dermatitis.
- The 10.7% relative improvement in uniqueness over the best TRIOMPHE variant supports the paper's claim that conditioning the generator directly on the latent transcriptomic state, rather than using the profile only to pick a source molecule, is what drives the gain.
Reading between the lines
- A clean control would condition the LSTM on a random 64-dimensional latent vector; if random-condition molecules match the target-profile molecules in Tanimoto similarity to known ligands, the expression profile is not causally steering generation.
- The disease case study uses 884-gene profiles while the VAE was trained on 978 LINCS genes; resolving this gene-set mapping and checking VAE reconstruction on disease profiles would make the disease results independently reproducible.
- Novelty is computed on canonical SMILES strings, so molecules sharing a scaffold but differing in a side chain both count as novel; a fingerprint- or scaffold-based novelty measure would give a stricter estimate of chemical-space exploration.
- The same conditioning scheme could be retrained on multi-cell-line and multi-dose response data; if it transfers, the approach would extend beyond the ten targets and three diseases reported here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces HVL2Mol, a two-stage generative model that combines a variational autoencoder (VAE) trained on 978-dimensional LINCS gene expression profiles with an LSTM-based SMILES generator. The VAE latent vector is intended to condition the LSTM during training and at inference; the authors test the model on ten target-protein perturbation profiles and on disease reversal profiles from CREEDS, reporting validity of 88.6%, uniqueness of 83.0%, novelty of 99.7%, and favorable maximum Tanimoto similarities to known ligands relative to ConGAN and TRIOMPHE. The central claim is that gene expression profiles can serve as a biological condition for de novo generation of hit-like molecules, and the paper includes a GitHub repository with source code.
Significance. If the conditioning mechanism were rigorously established, the paper would offer a simple and modular baseline for transcriptome-conditioned molecular generation, and the public code would be a useful resource. The main strengths are the use of a real transcriptomic dataset (LINCS MCF7 profiles), the comparison with two prior transcriptome-based methods, and the inclusion of a disease case study. However, the current evidence does not yet establish that the gene expression profile steers generation, because the headline Tanimoto metric is an order statistic over generated samples, the generator loss in Eq. (5) omits the conditioning variable, and no unconditioned or permutation controls are reported.
major comments (5)
- [Section 4.4 (Algorithm 1 and Figure 10)] The Tanimoto coefficients in Table 3 are maxima over the generated molecules for each target, as stated in Algorithm 1 (lines 15–16) and confirmed by the Figure 10 caption ("which have the highest Tanimoto coefficients"). A maximum over ~1,000 samples is an order statistic: it grows with sampling effort and can be large even when most generated molecules are unrelated to the target, so it does not measure whether the gene expression profile directs the generation. Please report distributional statistics (e.g., mean, median, and fraction above a threshold) over all valid generated molecules, and compare these statistics under identical sampling protocols for all methods.
- [Section 3.3, Eq. (5)] Equation (5) defines the generator loss as -Σ log p(y_i | X_{1:i-1}; φ), with no conditioning variable (z or FGx) in the argument. Although Section 3.3 states that the extracted features are concatenated with each SMILES token, the implemented conditional distribution is not specified and does not appear in the objective. Please state the exact conditional model (e.g., p(y_i | X_{1:i-1}, FGx; φ)) and update Eq. (5) accordingly; additionally, provide a control experiment with an unconditioned LSTM, a zero latent vector, or permuted expression profiles to show that the condition actually affects the output.
- [Section 4.1 (disease case study)] The VAE encoder is trained on 978-dimensional LINCS landmark-gene profiles, but the CREEDS disease profiles are described as containing 14,804 genes from which "the most relevant 884 genes" are extracted. The selection criterion for these 884 genes and the mapping from those genes to the 978 input dimensions are not described, so it is unclear how the disease reversal profiles can be fed to the model. Please specify the gene identifiers and the alignment procedure; without this, the disease case study is not reproducible.
- [Table 3 and Section 4.4] The known-ligand sets used to compute Tanimoto similarity are never defined: their source, version, size, and preprocessing (e.g., salt removal, canonicalization, filtering by target) are missing, which makes the comparison in Table 3 non-reproducible. All values in Table 3 and in the disease case study appear to be single-run results with no error bars; please report means and standard deviations over multiple random seeds, or state the number of runs and confirm stability.
- [Section 4.2 (validity)] The Validity bullet defines validity as "the ratio of valid molecules to the total number of training SMILES strings," which is inconsistent with the reported computation of 1171 valid molecules from the 1322 generated molecules in Section 4.3. The definition should refer to the generated set, and the text should be corrected to match the standard measure.
minor comments (8)
- [Abstract and title] The model is called HVL2Mol in the title and most of the paper, but the abstract uses HNN2Mol and HNNMol; please use one consistent name throughout.
- [Section 4.1] The GitHub link in Section 4.1 points to a repository named "Gx2Mol" while the paper describes "HVL2Mol"; please reconcile the names.
- [Section 4.4] In Section 4.4, "PIK3CCA" appears to be a typo for "PIK3CA".
- [Section 4.4] In Section 4.4, "STOA" should be "SOTA".
- [Section 4.2 (SA score)] The SA score description in Section 4.2 states that a higher SA score indicates greater ease of synthesis, which is the opposite of the standard interpretation of the Ertl–Schuffenhauer score; if a modified ease-of-synthesis score is used, define it explicitly.
- [Figure 10 caption] The Figure 10 caption refers to "the 1st column of the table," but Figure 10 is a figure; rephrase to refer to the figure's columns.
- [Section 4.2 (novelty)] The Novelty definition in Section 4.2 is phrased in a confusing way; clarify that novelty is the fraction of generated valid molecules whose canonical SMILES are absent from the training set.
- [Section 4.1 (model selection)] The description of model selection as "monitoring the convergence of the total loss" in Section 4.1 is vague; describe the early-stopping rule and the checkpoint selection procedure.
Circularity Check
The main structural-similarity benchmark is a maximum-Tanimoto selection made with the benchmark itself, so the hit-like comparison is partly forced by construction; the rest of the pipeline is not circular.
-
fitted input called prediction
[Section 4.4, Algorithm 1 (lines 15-16), and Fig. 10 caption]
"Calculate the Tanimoto coefficient using known ligands. Select the molecule with the maximum Tanimoto coefficient score as the candidate molecule. ... which have the highest Tanimoto coefficients with the corresponding known ligands."
The paper reports the generated molecule's Tanimoto coefficient as evidence that the model produces hit-like molecules, but the reported molecule is explicitly the sample with the maximum Tanimoto coefficient against the known-ligand set, selected using that same known-ligand set. A maximum over ~1000 generated samples is an upper order statistic: it is guaranteed to be at the tail of the similarity distribution even if the expression-profile condition has no effect on chemistry, and the paper provides no unconditioned or permuted-condition control showing that this maximum exceeds an unguided LSTM's maximum. The hit-like similarity value is therefore fitted to the benchmark by selection rather than predicted by the model, and the reported comparison partly compares selection artifacts.
full rationale
The derivation chain is not globally circular: the VAE and LSTM are trained on gene expression profiles and SMILES strings respectively, and the validity, uniqueness, novelty, QED, and SA metrics are computed on the generated molecules without fitting to the benchmark. I do not treat Eq. (5) omitting the conditioning variable, the undocumented mapping of the 884 CREEDS genes onto the 978-dimensional VAE input, or the absence of unconditioned controls as circularity; those are specification and generalizability gaps. The self-citations to TRIOMPHE and DRAGONET are ordinary baseline citations rather than load-bearing support: the knockdown/inhibitor correlation premise is also referenced to an external source [63], and no uniqueness theorem is imported. The single true circular step is the evaluation protocol: Table 3 and Fig. 10 select the maximum-Tanimoto molecule using the known-ligand benchmark and then cite that maximum as evidence of hit-likeness. This makes the structural-similarity 'prediction' an order statistic forced by construction, so the score is set to 6; this is partial circularity, not a fully circular derivation.
Assumptions & free parameters
free parameters (3)
- Number of selected genes for disease profiles =
884
- VAE latent dimension =
64
- Beta weight for KL divergence in VAE loss =
not stated
assumptions (4)
- domain assumption 978 LINCS landmark genes capture the drug-response signal needed for molecule generation
- domain assumption Knockdown and overexpression profiles of a target protein are sufficient surrogates for inhibitor and activator treatments
- domain assumption Multiplying a disease profile by -1 yields a therapeutic reversal profile
- standard math RDKit's validity, QED, SA, and ECFP4 Tanimoto are appropriate measures of hit-likeness
Cite this review
Pith. "Pith review of De Novo Generation of Hit-like Molecules from Gene Expression Profiles via Deep Learning." pith.science (2026). https://pith.science/paper/UCWRYSGH
@misc{pith2026241219422,
author = {Pith},
title = {Pith review of: De Novo Generation of Hit-like Molecules from Gene Expression Profiles via Deep Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/UCWRYSGH}},
note = {Machine review of arXiv:2412.19422}
}
read the original abstract
De novo generation of hit-like molecules is a challenging task in the drug discovery process. Most methods in previous studies learn the semantics and syntax of molecular structures by analyzing molecular graphs or simplified molecular input line entry system (SMILES) strings; however, they do not take into account the drug responses of the biological systems consisting of genes and proteins. In this study we propose a hybrid neural network, HNN2Mol, which utilizes gene expression profiles to generate molecular structures with desirable phenotypes for arbitrary target proteins. In the algorithm, a variational autoencoder is employed as a feature extractor to learn the latent feature distribution of the gene expression profiles. Then, a long short-term memory is leveraged as the chemical generator to produce syntactically valid SMILES strings that satisfy the feature conditions of the gene expression profile extracted by the feature extractor. Experimental results and case studies demonstrate that the proposed HNN2Mol model can produce new molecules with potential bioactivities and drug-like properties.
Figures
Figures from the paper (9 more)
Forward citations
Cited by 1 Pith paper
-
TextOmics-Guided Diffusion for Hit-like Molecular Generation
A joint omics-plus-text conditioning signal can generate chemically valid, novel hit-like molecules, demonstrated on the new TextOmics benchmark.
Reference graph
Works this paper leans on
-
[1]
Y . Ding, J. Tang, and F. Guo, “Identification of drug-targ et interac- tions via multi-view graph regularized link propagation mo del,” Neurocomputing, vol. 461, pp. 618–631, 2021
work page 2021
-
[2]
ChemoGraph: Interactive visual exploration o f the chemical space,
B. Kale, A. Clyde, M. Sun, A. Ramanathan, R. Stevens, and M. E. Papka, “ChemoGraph: Interactive visual exploration o f the chemical space,” Computer Graphics Forum, vol. 42, pp. 13–24, 2023
work page 2023
-
[3]
P . Ertl, “Cheminformatics analysis of organic substitu ents: identifi- cation of the most common substituents, calculation of subs tituent properties, and automatic identification of drug-like bioi sosteric groups,” Journal of chemical information and computer sciences , vol. 43, no. 2, pp. 374–380, 2003. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUS...
work page 2003
-
[4]
Molecular ge nerative graph neural networks for drug discovery ,
P . Bongini, M. Bianchini, and F. Scarselli, “Molecular ge nerative graph neural networks for drug discovery ,” Neurocomputing, vol. 450, pp. 242–252, 2021
work page 2021
-
[5]
Estimated resear ch and development investment needed to bring a new medicine to market, 2009-2018,
O. J. Wouters, M. McKee, and J. Luyten, “Estimated resear ch and development investment needed to bring a new medicine to market, 2009-2018,” Jama, vol. 323, no. 9, pp. 844–853, 2020
work page 2009
-
[6]
D. Chang, C.-H. Chen, and K. M. Lee, “A crowdsourcing deve l- opment approach based on a neuro-fuzzy network for creating innovative product concepts,” Neurocomputing, vol. 142, pp. 60– 72, 2014
work page 2014
-
[7]
A. Stecula, M. S. Hussain, and R. E. Viola, “Discovery of nov el inhibitors of a critical brain enzyme using a homology model and a deep convolutional neural network,” Journal of Medicinal Chemistry, vol. 63, no. 16, pp. 8867–8875, 2020
work page 2020
-
[8]
The light and dark sides of virtual scree ning: what is there to know?
A. Gimeno, M. J. Ojeda-Montes, S. Tom´ as-Hern´ andez, A. C ereto- Massagu´ e, R. Beltr´ an-Deb ´ on, M. Mulero, G. Pujadas, and S. Garcia-Vallv´ e, “The light and dark sides of virtual scree ning: what is there to know?” International journal of molecular sciences , vol. 20, no. 6, p. 1375, 2019
work page 2019
Show all 65 references
-
[9]
Machine lear ning in virtual screening,
J. L. Melville, E. K. Burke, and J. D. Hirst, “Machine lear ning in virtual screening,” Combinatorial chemistry & high throughput screening, vol. 12, no. 4, pp. 332–343, 2009
2009
-
[10]
Automated de novo drug desi gn: are we nearly there yet?
G. Schneider and D. E. Clark, “Automated de novo drug desi gn: are we nearly there yet?” Angewandte Chemie International Edition , vol. 58, no. 32, pp. 10 792–10 803, 2019
2019
-
[11]
A review on applications of com puta- tional methods in drug screening and design,
X. Lin, X. Li, and X. Lin, “A review on applications of com puta- tional methods in drug screening and design,” Molecules, vol. 25, no. 6, p. 1375, 2020
2020
-
[12]
Discovery and structure–activity analysis of selective e strogen receptor modulators via similarity-based virtual screeni ng,
J. Shen, J. Jiang, G. Kuang, C. Tan, G. Liu, J. Huang, and Y . Tang, “Discovery and structure–activity analysis of selective e strogen receptor modulators via similarity-based virtual screeni ng,” Euro- pean journal of medicinal chemistry , vol. 54, pp. 188–196, 2012
2012
-
[13]
Mol ec- ular graph enhanced transformer for retrosynthesis predic tion,
K. Mao, X. Xiao, T. Xu, Y . Rong, J. Huang, and P . Zhao, “Mol ec- ular graph enhanced transformer for retrosynthesis predic tion,” Neurocomputing, vol. 457, pp. 193–202, 2021
2021
-
[14]
Bifunctional tools to study adenosine receptors ,
C. Payne, J. K. Awalt, L. T. May , J. D. Tyndall, M. J ¨ org, a nd A. J. V ernall, “Bifunctional tools to study adenosine receptors ,” T opics in Medicinal Chemistry , pp. 1–43, 2022
2022
-
[15]
Molecula r property prediction and molecular design using a supervise d grammar variational autoencoder,
A. F. Oliveira, J. L. Da Silva, and M. G. Quiles, “Molecula r property prediction and molecular design using a supervise d grammar variational autoencoder,” Journal of Chemical Information and Modeling, vol. 62, no. 4, pp. 817–828, 2022
2022
-
[16]
Atte ntion- based generative models for de novo molecular design,
O. Dollar, N. Joshi, D. A. Beck, and J. Pfaendtner, “Atte ntion- based generative models for de novo molecular design,” Chemical Science, vol. 12, no. 24, pp. 8362–8372, 2021
2021
-
[17]
MolGAN: An implicit generative model for small molecular graphs. arxiv 2018,
N. De Cao and T. Kipf, “MolGAN: An implicit generative model for small molecular graphs. arxiv 2018,” arXiv preprint arXiv:1805.11973, 2019
2018 arXiv
-
[18]
Transf ormer- based objective-reinforced generative adversarial netwo rk to gen- erate desired molecules,
C. Li, C. Y amanaka, K. Kaitoh, and Y . Y amanishi, “Transf ormer- based objective-reinforced generative adversarial netwo rk to gen- erate desired molecules,” IJCAI, pp. 3884–3890, 2022
2022
-
[19]
SpotGAN: A reverse-transformer GAN generates scaffold-constrained molecules with property o ptimiza- tion,
C. Li and Y . Y amanishi, “SpotGAN: A reverse-transformer GAN generates scaffold-constrained molecules with property o ptimiza- tion,” Joint European Conference on Machine Learning and Knowledg e Discovery in Databases , pp. 323–338, 2023
2023
-
[20]
The impact of assay technology as applied to safety assessment in reducing compound attritio n in drug discovery ,
C. E. Thomas and Y . Will, “The impact of assay technology as applied to safety assessment in reducing compound attritio n in drug discovery ,” Expert Opinion on Drug Discovery , vol. 7, no. 2, pp. 109–122, 2012
2012
-
[21]
De novo structure-based drug design using deep learning,
S. R. Krishnan, N. Bung, S. R. Vangala, R. Srinivasan, G. Bul usu, and A. Roy , “De novo structure-based drug design using deep learning,” Journal of Chemical Information and Modeling , vol. 62, no. 21, pp. 5100–5109, 2021
2021
-
[22]
An in silico explaina ble multiparameter optimization approach for de novo drug desi gn against proteins from the central nervous system,
N. Bung, S. R. Krishnan, and A. Roy , “An in silico explaina ble multiparameter optimization approach for de novo drug desi gn against proteins from the central nervous system,” Journal of Chemical Information and Modeling , vol. 62, no. 11, pp. 2685–2695, 2022
2022
-
[23]
De novo generation of hit-like molecules from g ene expression signatures using artificial intelligence,
O. M´ endez-Lucio, B. Baillif, D.-A. Clevert, D. Rouqui ´ e, and J. Wichard, “De novo generation of hit-like molecules from g ene expression signatures using artificial intelligence,” Nature commu- nications, vol. 11, no. 1, p. 10, 2020
2020
-
[24]
TRIOMPHE: Transcriptome- based inference and generation of molecules with desired phenoty pes by machine learning,
K. Kaitoh and Y . Y amanishi, “TRIOMPHE: Transcriptome- based inference and generation of molecules with desired phenoty pes by machine learning,” Journal of Chemical Information and Modeling , vol. 61, no. 9, pp. 4303–4320, 2021
2021
-
[25]
Captur ing temporal dynamics of users’ preferences from purchase hist ory big data for recommendation system,
C. Li, M. He, M. Qaosar, S. Ahmed, and Y . Morimoto, “Captur ing temporal dynamics of users’ preferences from purchase hist ory big data for recommendation system,” 2018 IEEE International Conference on Big Data (Big Data) , pp. 5372–5374, 2018
2018
-
[26]
A multi-factor approa ch for stock price prediction by using recurrent neural networ ks,
X. Zhang, C. Li, and Y . Morimoto, “A multi-factor approa ch for stock price prediction by using recurrent neural networ ks,” Bulletin of networking, computing, systems, and software , vol. 8, no. 1, pp. 9–13, 2019
2019
-
[27]
Time series forecasting of petro leum production using deep LSTM recurrent networks,
A. Sagheer and M. Kotb, “Time series forecasting of petro leum production using deep LSTM recurrent networks,” Neurocomput- ing, vol. 323, pp. 203–213, 2019
2019
-
[28]
From theory to experiment: transformer-based generation enables rapid discovery of novel reactions,
X. Wang, C. Y ao, Y . Zhang, J. Y u, H. Qiao, C. Zhang, Y . Wu, R. Bai, and H. Duan, “From theory to experiment: transformer-based generation enables rapid discovery of novel reactions,” Journal of Cheminformatics, vol. 14, no. 1, pp. 1–14, 2022
2022
-
[29]
C. G. Wermuth, The practice of medicinal chemistry. Academic Press, 2011
2011
-
[30]
Structure-based design , synthesis, and evaluation of peptide-mimetic SARS 3CL prote ase inhibitors,
K. Akaji, H. Konno, H. Mitsui, K. Teruya, Y . Shimamoto, Y . Hattori, T. Ozaki, M. Kusunoki, and A. Sanjoh, “Structure-based design , synthesis, and evaluation of peptide-mimetic SARS 3CL prote ase inhibitors,” Journal of medicinal chemistry , vol. 54, no. 23, pp. 7962– 7973, 2011
2011
-
[31]
Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation,
M. Krenn, F. H¨ ase, A. Nigam, P . Friederich, and A. Aspur u- Guzik, “Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation,” Machine Learning: Science and T echnology, vol. 1, no. 4, p. 045024, 2020
2020
-
[32]
Junction tree var iational au- toencoder for molecular graph generation,
W. Jin, R. Barzilay , and T. Jaakkola, “Junction tree var iational au- toencoder for molecular graph generation,” International conference on machine learning , pp. 2323–2332, 2018
2018
-
[33]
Interpretable molec ular graph generation via monotonic constraints,
Y . Du, X. Guo, A. Shehu, and L. Zhao, “Interpretable molec ular graph generation via monotonic constraints,” Proceedings of the 2022 SIAM International Conference on Data Mining (SDM) , pp. 73– 81, 2022
2022
-
[34]
Small molecul e generation via disentangled representation learning,
Y . Du, X. Guo, Y . Wang, A. Shehu, and L. Zhao, “Small molecul e generation via disentangled representation learning,” Bioinformat- ics, vol. 38, no. 12, pp. 3200–3208, 2022
2022
-
[35]
MoFlow: an invertible flow model for gen- erating molecular graphs,
C. Zang and F. Wang, “MoFlow: an invertible flow model for gen- erating molecular graphs,” Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data min ing, pp. 617–626, 2020
2020
-
[36]
FastFlows: Flow- based models for molecular graph generation,
N. C. Frey , V . Gadepally , and B. Ramsundar, “FastFlows: Flow- based models for molecular graph generation,” arXiv preprint arXiv:2201.12419, 2022
2022 arXiv
-
[37]
DiGress: Discrete denoising diffusion for gr aph gen- eration,
C. Vignac, I. Krawczuk, A. Siraudin, B. Wang, V . Cevher, a nd P . Frossard, “DiGress: Discrete denoising diffusion for gr aph gen- eration,” Proceedings of the 11th International Conference on Learni ng Representations, 2023
2023
-
[38]
Continuous control with deep rei n- forcement learning,
T. P . Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez , Y . Tassa, D. Silver, and D. Wierstra, “Continuous control with deep rei n- forcement learning,” arXiv preprint arXiv:1509.02971 , 2015
2015 arXiv
-
[39]
Adversarial learned mol ecular graph inference and generation,
S. P ¨ olsterl and C. Wachinger, “Adversarial learned mol ecular graph inference and generation,” Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 202 0, Ghent, Belgium, September 14–18, 2020, Proceedings, Part I I, pp. 173– 189, 2021
2020
-
[40]
Gr ammar variational autoencoder,
M. J. Kusner, B. Paige, and J. M. Hern´ andez-Lobato, “Gr ammar variational autoencoder,” International conference on machine learn- ing, pp. 1945–1954, 2017
1945
-
[41]
DNMG: Deep molecular generative model by fusion of 3d information for de novo drug design,
T. Song, Y . Ren, S. Wang, P . Han, L. Wang, X. Li, and A. Rodriguez- Pat ´ on, “DNMG: Deep molecular generative model by fusion of 3d information for de novo drug design,” Methods, vol. 211, pp. 10–22, 2023
2023
-
[42]
Monte-carlo simulation balan cing,
D. Silver and G. Tesauro, “Monte-carlo simulation balan cing,” Proceedings of the 26th Annual International Conference on Machine Learning, pp. 945–952, 2009
2009
-
[43]
A review of molecular representation in the age of machine learning,
D. S. Wigh, J. M. Goodman, and A. A. Lapkin, “A review of molecular representation in the age of machine learning,” Wiley Interdisciplinary Reviews: Computational Molecular Scie nce, vol. 12, no. 5, p. e1603, 2022
2022
-
[44]
SELF- IES and the future of molecular string representations,
M. Krenn, Q. Ai, S. Barthel, N. Carson, A. Frei, N. C. Frey , P . Friederich, T. Gaudin, A. A. Gayle, K. M. Jablonka et al. , “SELF- IES and the future of molecular string representations,” Patterns, vol. 3, no. 10, 2022
2022
-
[45]
TSI S: A supplementary algorithm to t-SMILES for fragment-based mol ec- ular representation,
J.-N. Wu, T. Wang, L.-J. Tang, H.-L. Wu, and R.-Q. Y u, “TSI S: A supplementary algorithm to t-SMILES for fragment-based mol ec- ular representation,” arXiv preprint arXiv:2402.02164 , 2024. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 12
2024 arXiv
-
[46]
Predicting physiolo gical effects of chemical substances using natural language proc essing,
S. Mukherjee, J. Ben-Joseph, M. Campos, P . Malla, H. Nguy en, A. Pham, T. Oates, and V . Janarthanan, “Predicting physiolo gical effects of chemical substances using natural language proc essing,” 2021 IEEE Canadian Conference on Electrical and Computer En gineer- ing (CCECE)...
2021
-
[47]
Advances in machine lea rn- ing with chemical language models in molecular property and reaction outcome predictions,
M. Das, A. Ghosh, and R. B. Sunoj, “Advances in machine lea rn- ing with chemical language models in molecular property and reaction outcome predictions,” Journal of Computational Chemistry , 2024
2024
-
[48]
Exploring chemical space — generative models and their evaluation,
M. V ogt, “Exploring chemical space — generative models and their evaluation,” Artificial Intelligence in the Life Sciences , vol. 3, p. 100064, 2023
2023
-
[49]
Lifelong genera- tive modeling,
J. Ramapuram, M. Gregorova, and A. Kalousis, “Lifelong genera- tive modeling,” Neurocomputing, vol. 404, pp. 381–400, 2020
2020
-
[50]
Kullback-leibler divergence,
J. M. Joyce, “Kullback-leibler divergence,” International encyclope- dia of statistical science , pp. 720–722, 2011
2011
-
[51]
LINCS canvas browser: interactive web app to query , browse and interrogate lincs l1000 gene expression signatures,
Q. Duan, C. Flynn, M. Niepel, M. Hafner, J. L. Muhlich, N. F. Fernandez, A. D. Rouillard, C. M. Tan, E. Y . Chen, T. R. Golubet al., “LINCS canvas browser: interactive web app to query , browse and interrogate lincs l1000 gene expression signatures,” Nucleic acids research, vo...
2014
-
[52]
Extraction and analysis of signatures from the gene expression omnibus by the crowd,
Z. Wang, C. D. Monteiro, K. M. Jagodnik, N. F. Fernandez, G. W. Gundersen, A. D. Rouillard, S. L. Jenkins, A. S. Feldmann, K. S. Hu, M. G. McDermott et al., “Extraction and analysis of signatures from the gene expression omnibus by the crowd,” Nature commu- nications, vol. 7, ...
2016
-
[53]
Dropout: a simple way to prevent neural networks from overfitting,
N. Srivastava, G. Hinton, A. Krizhevsky , I. Sutskever, an d R. Salakhutdinov , “Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research , vol. 15, no. 1, pp. 1929–1958, 2014
1929
-
[54]
Adam: A method for stochastic opt imiza- tion,
D. P . Kingma and J. Ba, “Adam: A method for stochastic opt imiza- tion,” arXiv preprint arXiv:1412.6980 , 2014
2014 arXiv
-
[55]
Quantifying the chemical beauty of drugs,
G. R. Bickerton, G. V . Paolini, J. Besnard, S. Muresan, an d A. L. Hopkins, “Quantifying the chemical beauty of drugs,” Nature chemistry, vol. 4, no. 2, pp. 90–98, 2012
2012
-
[56]
Estimation of synthetic a ccessibility score of drug-like molecules based on molecular complexity and fragment contributions,
P . Ertl and A. Schuffenhauer, “Estimation of synthetic a ccessibility score of drug-like molecules based on molecular complexity and fragment contributions,” Journal of cheminformatics , vol. 1, no. 1, pp. 1–11, 2009
2009
-
[57]
Life beyond the tanimoto co- efficient: similarity measures for interaction fingerprint s,
A. R ´ acz, D. Bajusz, and K. H´ eberger, “Life beyond the tanimoto co- efficient: similarity measures for interaction fingerprint s,” Journal of cheminformatics, vol. 10, no. 1, pp. 1–12, 2018
2018
-
[58]
Rdkit documentation,
G. Landrum, “Rdkit documentation,” Release, vol. 1, no. 1-79, p. 4, 2013
2013
-
[59]
MolFinder: an evolutionary algorit hm for the global optimization of molecular properties and the ext ensive exploration of chemical space using smiles,
Y . Kwon and J. Lee, “MolFinder: an evolutionary algorit hm for the global optimization of molecular properties and the ext ensive exploration of chemical space using smiles,” Journal of cheminfor- matics, vol. 13, pp. 1–14, 2021
2021
-
[60]
Prediction of drug-likeness using grap h convolutional attention network,
J. Sun, M. Wen, H. Wang, Y . Ruan, Q. Y ang, X. Kang, H. Zhang, Z. Zhang, and H. Lu, “Prediction of drug-likeness using grap h convolutional attention network,” Bioinformatics, vol. 38, no. 23, pp. 5262–5269, 2022
2022
-
[61]
Extended-connectivity fingerpr ints,
D. Rogers and M. Hahn, “Extended-connectivity fingerpr ints,” Journal of chemical information and modeling , vol. 50, no. 5, pp. 742– 754, 2010
2010
-
[62]
Improving MRI segmentation with probabilistic GHSOM and multiobjective optimization,
A. Ortiz, J. M. Gorriz, J. Ram´ ırez, D. Salas-Gonzalez, A . D. N. Initiative et al. , “Improving MRI segmentation with probabilistic GHSOM and multiobjective optimization,” Neurocomputing, vol. 114, pp. 118–131, 2013
2013
-
[63]
Single-layer artificial neural networks for gene expressio n analy- sis,
A. Narayanan, E. C. Keedwell, J. Gamalielsson, and S. Tat ineni, “Single-layer artificial neural networks for gene expressio n analy- sis,” Neurocomputing, vol. 61, pp. 217–240, 2004
2004
-
[64]
De novo drug design based on patient gene expression profiles vi a deep learning,
C. Y amanaka, S. Uki, K. Kaitoh, M. Iwata, and Y . Y amanishi , “De novo drug design based on patient gene expression profiles vi a deep learning,” Molecular Informatics , vol. 42, no. 8-9, p. 2300064, 2023
2023
-
[65]
Prediction of cancer drugs by chemical-chem ical interactions,
J. Lu, G. Huang, H.-P . Li, K.-Y . Feng, L. Chen, M.-Y . Zhen g, and Y .-D. Cai, “Prediction of cancer drugs by chemical-chem ical interactions,” PLoS One, vol. 9, no. 2, p. e87791, 2014. Chen Li received the PhD degree in engineering from Hiroshima University , Japan, in 2019...
2014
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.