Fetching the paper…
Reading the bibliography…
Transformer-based models trained on large and general purpose datasets consisting of molecular strings have recently emerged as a powerful tool for successfully modeling various structure-property relations.
Scaling laws for neural language models (2020)
Kaplan, J. et al · 2001
Earlier work this paper cites.
ZINC–a free database of commercially available compounds for virtual screening
Irwin, J. J. & Shoichet, B. K · 2005
Earlier work this paper cites.
Extracting training data from large language models (2021)
Carlini, N. et al · 2012
Earlier work this paper cites.
Matched molecular pair analysis in drug discovery
Dossetter, A. G., Griffen, E. J. & Leach, A. G · 2013
Earlier work this paper cites.
Pubchem substance and compound databases
Kim, S. et al · 2016
Earlier work this paper cites.
Molecular de-novo design through deep reinforcement learning
Olivecrona, M., Blaschke, T., Engkvist, O. & Chen, H · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A. et al · 2017
Earlier work this paper cites.
Grammar variational autoencoder
Kusner, M. J., Paige, B. & Hernández-Lobato, J. M · 2017
Earlier work this paper cites.
Molecular sets (MOSES): A benchmarking platform for molecular generation models
Polykovskiy, D. et al · 2018
Earlier work this paper cites.
Junction tree variational autoencoder for molecular graph generation
Jin, W., Barzilay, R. & Jaakkola, T · 2018
Earlier work this paper cites.
Fréchet ChemNet distance: a metric for generative models for molecules in drug discovery
Preuer, K., Renz, P., Unterthiner, T., Hochreiter, S. & Klambauer, G · 2018
Earlier work this paper cites.
Graph convolutional policy network for goal-directed molecular graph generation
You, J., Liu, B., Ying, Z., Pande, V. & Leskovec, J · 2018
Earlier work this paper cites.
Automatic chemical design using a data-driven continuous representation of molecules
Gómez-Bombarelli, R. et al · 2018
Earlier work this paper cites.
PubChem 2019 update: improved access to chemical data
Kim, S. et al · 2018
Earlier work this paper cites.
Syntax-directed variational autoencoder for structured data
Dai, H., Tian, Y., Dai, B., Skiena, S. & Song, L · 2018
Earlier work this paper cites.
Deepsmiles: an adaptation of smiles for use in machine-learning of chemical structures
O’Boyle, N. & Dalke, A · 2018
Earlier work this paper cites.
Constrained generation of semantically valid graphs via regularizing variational autoencoders
Ma, T., Chen, J. & Xiao, C · 2018
Earlier work this paper cites.
MolGAN: An implicit generative model for small molecular graphs
De Cao, N. & Kipf, T · 2018
Earlier work this paper cites.
Anthropogenic biases in chemical reaction data hinder exploratory inorganic synthesis
Jia, X. et al · 2019
Earlier work this paper cites.
Guacamol: benchmarking models for de novo molecular design
Brown, N., Fiscato, M., Segler, M. H. & Vaucher, A. C · 2019
Earlier work this paper cites.
Optimization of molecules via deep reinforcement learning
Zhou, Z., Kearnes, S., Li, L., Zare, R. N. & Riley, P · 2019
Earlier work this paper cites.
Improved precision and recall metric for assessing generative models
Kynkäänniemi, T., Karras, T., Laine, S., Lehtinen, J. & Aila, T · 2019
Cited alongside, same era.
Molecular transformer: A model for uncertainty-calibrated chemical reaction prediction
Schwaller, P. et al · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K · 2019
Cited alongside, same era.
Learning multimodal graph-to-graph translation for molecular optimization (2019)
Jin, W., Yang, K., Barzilay, R. & Jaakkola, T · 2019
Cited alongside, same era.
Deep learning enables rapid identification of potent DDR1 kinase inhibitors
Zhavoronkov, A. et al · 2019
Cited alongside, same era.
Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation
Deduplicating training data makes language models better (2022)
Lee, K. et al · 2022
Later among the works it cites.
Large-scale chemical language representations capture molecular structure and properties
Ross, J. et al · 2022
Later among the works it cites.
Limo: Latent inceptionism for targeted molecule generation
Eckmann, P. et al · 2022
Later among the works it cites.
Chemformer: a pre-trained transformer for computational chemistry
Irwin, R., Dimitriadis, S., He, J. & Bjerrum, E. J · 2022
Later among the works it cites.
Artificial intelligence foundation for therapeutic science
Huang, K. et al · 2022
Later among the works it cites.
Augmenting molecular deep generative models with topological data analysis representations
Schiff, Y., Chenthamarakshan, V., Hoffman, S. C., Ramamurthy, K. N. & Das, P · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Krenn, M., Häse, F., Nigam, A., Friederich, P. & Aspuru-Guzik, A · 2020
Cited alongside, same era.
Smiles-based deep generative scaffold decorator for de-novo drug design
Arús-Pous, J. et al · 2020
Cited alongside, same era.
Mol-cyclegan: a generative model for molecular optimization
Maziarka, Ł. et al · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N. & Fleuret, F · 2020
Cited alongside, same era.
Cogmol: Target-specific and selective drug design for covid-19 using deep generative models
Chenthamarakshan, V. et al · 2020
Cited alongside, same era.
Hierarchical generation of molecular graphs using structural motifs
Jin, W., Barzilay, R. & Jaakkola, T · 2020
Cited alongside, same era.
Scaling laws for neural machine translation (2021)
Ghorbani, B. et al · 2021
Cited alongside, same era.
Later among the works it cites.
Perplexity-based molecule ranking and bias estimation of chemical language models
Moret, M., Grisoni, F., Katzberger, P. & Schneider, G · 2022
Later among the works it cites.
Quantifying memorization across neural language models (2023)
Carlini, N. et al · 2023
Later among the works it cites.
Large language models struggle to learn long-tail knowledge (2023)
Kandpal, N., Deng, H., Roberts, A., Wallace, E. & Raffel, C · 2023
Later among the works it cites.
Tartarus: A benchmarking platform for realistic and practical inverse molecular design (2023)
Nigam, A. et al · 2023
Later among the works it cites.
Druglike molecule datasets for drug discovery, DOI: 10.5281/zenodo.7547717 (2023)
Lee, J · 2023
Later among the works it cites.
Matched molecular pair analysis in drug discovery: methods and recent applications
Yang, Z. et al · 2023
Later among the works it cites.
Gargoyles: An open source graph-based molecular optimization method based on deep reinforcement learning
Erikawa, D., Yasuo, N., Suzuki, T., Nakamura, S. & Sekijima, M · 2023
Later among the works it cites.
Neural scaling of deep chemical models
Frey, N. C. et al · 2023
Later among the works it cites.
Uncorrupt smiles: a novel approach to de novo design
Schoenmaker, L., Béquignon, O. J., Jespers, W. & van Westen, G. J · 2023
Later among the works it cites.
Group selfies: a robust fragment-based molecular string representation
Cheng, A. H. et al · 2023
Later among the works it cites.
Calibrated language models must hallucinate (2024)
Kalai, A. T. & Vempala, S. S · 2024
Closest in time.
Large language monkeys: Scaling inference compute with repeated sampling (2024)
Brown, B. et al · 2024
Closest in time.
Domain-agnostic molecular generation with chemical feedback (2024)
Fang, Y. et al · 2024
Closest in time.
Invalid smiles are beneficial rather than detrimental to chemical language models
Skinnider, M. A · 2024
Closest in time.