Fetching the paper…
Reading the bibliography…
We discover a robust self-supervised strategy tailored towards molecular representations for generative masked language models through a series of tailored, in-depth ablations.
The generation of a unique machine description for chemical structures-a technique developed at chemical abstracts service
Harry L Morgan · 1965
Earlier work this paper cites.
The properties of known drugs. 1. molecular frameworks
Guy W Bemis and Mark A Murcko · 1996
Earlier work this paper cites.
Review of organic functional groups: introduction to medicinal organic chemistry
Thomas L Lemke · 2003
Earlier work this paper cites.
Circular fingerprints: flexible molecular descriptors with applications from physical chemistry to adme
Robert C Glen, Andreas Bender, Catrin H Arnby, Lars Carlsson, Scott Boyer, and James Smith · 2006
Earlier work this paper cites.
Better fine-tuning by reducing representational collapse
Armen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal, Luke Zettlemoyer, and Sonal Gupta · 2008
Earlier work this paper cites.
Benchmark data set for in silico prediction of ames mutagenicity
Katja Hansen, Sebastian Mika, Timon Schroeter, Andreas Sutter, Antonius ter Laak, Thomas Steger-Hartmann, Nikolaus Heinrich, and Klaus-Robert Müller · 2009
Earlier work this paper cites.
Extended-connectivity fingerprints
David Rogers and Mathew Hahn · 2010
Earlier work this paper cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta · 2012
Earlier work this paper cites.
Extraction of chemical structures and reactions from the literature
Daniel Mark Lowe · 2012
Earlier work this paper cites.
A bayesian approach to in silico blood-brain barrier penetration modeling
Ines Filipa Martins, Ana L Teixeira, Luis Pinheiro, and Andre O Falcao · 2012
Earlier work this paper cites.
Rdkit: A software suite for cheminformatics, computational chemistry, and predictive modeling, 2013
Greg Landrum et al · 2013
Earlier work this paper cites.
Multi-task neural networks for qsar predictions
George E Dahl, Navdeep Jaitly, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Convolutional networks on graphs for learning molecular fingerprints
David K Duvenaud, Dougal Maclaurin, Jorge Iparraguirre, Rafael Bombarell, Timothy Hirzel, Alán Aspuru-Guzik, and Ryan P Adams · 2015
Earlier work this paper cites.
The history and development of quantitative structure-activity relationships (qsars)
John Dearden · 2016
Earlier work this paper cites.
Molecular graph convolutions: moving beyond fingerprints
Steven Kearnes, Kevin McCloskey, Marc Berndl, Vijay Pande, and Patrick Riley · 2016
Earlier work this paper cites.
Scikit-learn
Oliver Kramer · 2016
Earlier work this paper cites.
Mutagenic and carcinogenic structural alerts and their mechanisms of action
Alja Plošnik, Marjan Vračko, and Marija Sollner Dolenc · 2016
Earlier work this paper cites.
Predicting organic reaction outcomes with weisfeiler-lehman network
Wengong Jin, Connor Coley, Regina Barzilay, and Tommi Jaakkola · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
A generative model for electron paths
John Bradshaw, Matt J Kusner, Brooks Paige, Marwin HS Segler, and José Miguel Hernández-Lobato · 2018
Earlier work this paper cites.
Retrosynthesis: Computer says yes
Stephen G Davey · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
In silico prediction of chemical genotoxicity using machine learning methods and structural alerts
Defang Fan, Hongbin Yang, Fuxing Li, Lixia Sun, Peiwen Di, Weihua Li, Yun Tang, and Guixia Liu · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Mol2vec: unsupervised machine learning approach with chemical intuition
Sabrina Jaeger, Simone Fulle, and Samo Turk · 2018
Cited alongside, same era.
Machine learning with physicochemical relationships: solubility prediction in organic solvents and water
Samuel Boobier, David RJ Hose, A John Blacker, and Bao N Nguyen · 2020
Later among the works it cites.
Chemberta: Large-scale self-supervised pretraining for molecular property prediction
Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar · 2020
Later among the works it cites.
Unsupervised cross-lingual representation learning for speech recognition
Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdelrahman Mohamed, and Michael Auli · 2020
Later among the works it cites.
Strategies for pre-training graph neural networks
Weihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik, Percy Liang, Vijay Pande, and Jure Leskovec · 2020
Later among the works it cites.
Zinc20—a free ultralarge-scale chemical database for ligand discovery
John J Irwin, Khanh G Tang, Jennifer Young, Chinzorig Dandarchuluun, Benjamin R Wong, Munkhzul Khurelbaatar, Yurii S Moroz, John Mayfield, and Roger A Sayle · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
Fréchet chemnet distance: a metric for generative models for molecules in drug discovery
Kristina Preuer, Philipp Renz, Thomas Unterthiner, Sepp Hochreiter, and Gunter Klambauer · 2018
Cited alongside, same era.
“found in translation”: predicting outcomes of complex organic chemistry reactions using neural sequence-to-sequence models
Philippe Schwaller, Theophile Gaudin, David Lanyi, Costas Bekas, and Teodoro Laino · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Cited alongside, same era.
Moleculenet: a benchmark for molecular machine learning
Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande · 2018
Cited alongside, same era.
A graph-convolutional neural network model for the prediction of chemical reactivity
Connor W Coley, Wengong Jin, Luke Rogers, Timothy F Jamison, Tommi S Jaakkola, William H Green, Regina Barzilay, and Klavs F Jensen · 2019
Cited alongside, same era.
Graph transformation policy network for chemical reaction prediction
Kien Do, Truyen Tran, and Svetha Venkatesh · 2019
Cited alongside, same era.
Later among the works it cites.
Captum: A unified and generic model interpretability library for pytorch, 2020
Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz-Richardson · 2020
Later among the works it cites.
Pre-training via paraphrasing, 2020
Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida Wang, and Luke Zettlemoyer · 2020
Later among the works it cites.
Self-supervised graph transformer on large-scale molecular data
Yu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie, Ying Wei, Wenbing Huang, and Junzhou Huang · 2020
Later among the works it cites.
Learning graph models for template-free retrosynthesis
Vignesh Ram Somnath, Charlotte Bunne, Connor W Coley, Andreas Krause, and Regina Barzilay · 2020
Later among the works it cites.
Energy-based view of retrosynthesis
Ruoxi Sun, Hanjun Dai, Li Li, Steven Kearnes, and Bo Dai · 2020
Later among the works it cites.
State-of-the-art augmented nlp transformer models for direct and single-step retrosynthesis
Igor V Tetko, Pavel Karpov, Ruud Van Deursen, and Guillaume Godin · 2020
Later among the works it cites.
Htlm: Hyper-text pre-training and prompting of language models
Armen Aghajanyan, Dmytro Okhonko, Mike Lewis, Mandar Joshi, Hu Xu, Gargi Ghosh, and Luke Zettlemoyer · 2021
Later among the works it cites.
Fairscale: A general purpose modular pytorch library for high performance and large scale training
Mandeep Baines, Shruti Bhosale, Vittorio Caggiano, Naman Goyal, Siddharth Goyal, Myle Ott, Benjamin Lefaudeux, Vitaliy Liptchinsky, Mike Rabbat, Sam Sheiffer, Anjali Sridhar, and Min Xu · 2021
Later among the works it cites.
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, and Furu Wei · 2021
Later among the works it cites.
Could graph neural networks learn better molecular representation for drug discovery? a comparison study of descriptor-based and graph-based models
Dejun Jiang, Zhenxing Wu, Chang-Yu Hsieh, Guangyong Chen, Ben Liao, Zhe Wang, Chao Shen, Dongsheng Cao, Jian Wu, and Tingjun Hou · 2021
Later among the works it cites.
Do large scale molecular language representations capture important structural information?
Jerret Ross, Brian Belgodere, Vijil Chenthamarakshan, Inkit Padhi, Youssef Mroueh, and Payel Das · 2021
Later among the works it cites.
Retrocomposer: Discovering novel reactions by composing templates for retrosynthesis prediction
Chaochao Yan, Peilin Zhao, Chan Lu, Yang Yu, and Junzhou Huang · 2021
Later among the works it cites.
Cm3: A causal masked multimodal model of the internet
Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, et al · 2022
Closest in time.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Closest in time.
Chemformer: a pre-trained transformer for computational chemistry
Ross Irwin, Spyridon Dimitriadis, Jiazhen He, and Esben Jannik Bjerrum · 2022
Closest in time.
3d infomax improves gnns for molecular property prediction
Hannes Stärk, Dominique Beaini, Gabriele Corso, Prudencio Tossou, Christian Dallago, Stephan Günnemann, and Pietro Liò · 2022
Closest in time.