Fetching the paper…
Reading the bibliography…
We show that large language model (LLMs) can be transformed via supervised fine-tuning (SFT) of engineered prompts into SmileyLlama for exploring the chemical space of drug molecules.
1909
Earlier work this paper cites.
Weininger, D. SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Comput. Sci. 1988
1988
Earlier work this paper cites.
Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Computation 1997
1997
Earlier work this paper cites.
Lipinski, C. A.; Lombardo, F.; Dominy, B. W.; Feeney, P. J. Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings1PII of original article: S0169-409X(96)00423-1. The article was originally published in Advanced Drug Delivery Reviews 23 (1997) 3–25.1. Advanced Drug Delivery Reviews 2001
1998
Earlier work this paper cites.
Rosenfeld, R. Two decades of statistical language modeling: where do we go from here? Proceedings of the IEEE 2000
2000
Earlier work this paper cites.
2002
Earlier work this paper cites.
Brown, T. B. et al. Language Models are Few-Shot Learners. http://arxiv.org/abs/2005.14165
2005
Earlier work this paper cites.
Hunter, J. D. Matplotlib: A 2D graphics environment. Computing in Science & Engineering 2007
2007
Earlier work this paper cites.
Landrum, G. RDKit: Open-Source Cheminformatics Software. 2016
2016
Earlier work this paper cites.
Gupta, A.; Müller, A. T.; Huisman, B. J. H.; Fuchs, J. A.; Schneider, P.; Schneider, G. Generative Recurrent Networks for De Novo Drug Design. Molecular Informatics 2017
2017
Earlier work this paper cites.
Cao, N. D.; Kipf, T. MolGAN: An implicit generative model for small molecular graphs. ArXiv 2018
2018
Earlier work this paper cites.
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I. Improving Language Understanding by Generative Pre-Training. 2018; https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf
2018
Earlier work this paper cites.
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I. Language Models are Unsupervised Multitask Learners. 2018
2018
Earlier work this paper cites.
Preuer, K.; Renz, P.; Unterthiner, T.; Hochreiter, S.; Klambauer, G. Fréchet ChemNet distance: a metric for generative models for molecules in drug discovery. J. Chem. Inform. Model. 2018
2018
Earlier work this paper cites.
Krenn, M.; Hase, F.; Nigam, A.; Friederich, P.; Aspuru-Guzik, A. Self-referencing embedded strings (SELFIES): A 100
2019
Earlier work this paper cites.
Zhou, Z.; Kearnes, S.; Li, L.; Zare, R. N.; Riley, P. Optimization of Molecules via Deep Reinforcement Learning. Scientific Reports 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Brown, N.; Fiscato, M.; Segler, M. H.; Vaucher, A. C. GuacaMol: Benchmarking Models for de Novo Molecular Design. Journal of Chemical Information and Modeling 2019
2019
Earlier work this paper cites.
Blaschke, T.; Arús-Pous, J.; Chen, H.; Margreitter, C.; Tyrchan, C.; Engkvist, O.; Papadopoulos, K.; Patronov, A. REINVENT 2.0: An AI Tool for De Novo Drug Design. Journal of chemical information and modeling 2020
2020
Earlier work this paper cites.
Rasley, J.; Rajbhandari, S.; Ruwase, O.; He, Y. DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters. Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2020
2020
Earlier work this paper cites.
Zhang, M.; He, Y. Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping. Advances in Neural Information Processing Systems. 2020
2020
Earlier work this paper cites.
Tong, X.; Liu, X.; Tan, X.; Li, X.; Jiang, J.; Xiong, Z.; Xu, T.; Jiang, H.; Qiao, N.; Zheng, M. Generative Models for De Novo Drug Design. Journal of medicinal chemistry 2021
2021
Earlier work this paper cites.
Flam-Shepherd, D.; Zhu, K.; Aspuru-Guzik, A. Language models can learn complex molecular distributions. Nature Communications 2021
2021
Earlier work this paper cites.
Skinnider, M.; Stacey, R.; Wishart, D.; Foster, L. Chemical language models enable navigation in sparsely populated chemical space. Nature Machine Intelligence 2021
2021
Earlier work this paper cites.
Huang, K.; Fu, T.; Gao, W.; Zhao, Y.; Roohani, Y.; Leskovec, J.; Coley, C. W.; Xiao, C.; Sun, J.; Zitnik, M. Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and Development. Proceedings of Neural Information Processing Systems, NeurIPS Datasets and Benchmarks 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Tang, H.; Gan, S.; Awan, A. A.; Rajbhandari, S.; Li, C.; Lian, X.; Liu, J.; Zhang, C.; He, Y. 1-bit Adam: Communication Efficient Large-Scale Training with Adam’s Convergence Speed. International Conference on Machine Learning. 2021
2021
Earlier work this paper cites.
2021
Cited alongside, same era.
2022
Cited alongside, same era.
Huang, K.; Fu, T.; Gao, W.; Zhao, Y.; Roohani, Y.; Leskovec, J.; Coley, C. W.; Xiao, C.; Sun, J.; Zitnik, M. Artificial intelligence foundation for therapeutic science. Nature Chemical Biology 2022
2022
Cited alongside, same era.
Gao, W.; Fu, T.; Sun, J.; Coley, C. Sample efficiency matters: a benchmark for practical molecular optimization. Advances in Neural Information Processing Systems 2022
2022
Cited alongside, same era.
Anthony, Q.; Awan, A. A.; Rasley, J.; He, Y.; Shafi, A.; Abduljabbar, M.; Subramoni, H.; Panda, D. MCR-DL: Mix-and-Match Communication Runtime for Deep Learning. IEEE International Parallel and Distributed Processing Symposium. 2023
2023
Later among the works it cites.
Singh, S.; Ruwase, O.; Awan, A. A.; Rajbhandari, S.; He, Y.; Bhatele, A. A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training. Proceedings of the 37th International Conference on Supercomputing. 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dao, T.; Fu, D. Y.; Ermon, S.; Rudra, A.; Ré, C. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. Advances in Neural Information Processing Systems (NeurIPS). 2022
2022
Cited alongside, same era.
Gugger, S.; Debut, L.; Wolf, T.; Schmid, P.; Mueller, Z.; Mangrulkar, S.; Sun, M.; Bossan, B. Accelerate: Training and inference at scale made simple, efficient and adaptable. https://github.com/huggingface/accelerate
2022
Cited alongside, same era.
Li, C.; Awan, A. A.; Tang, H.; Rajbhandari, S.; He, Y. 1-bit LAMB: Communication Efficient Large-Scale Large-Batch Training with LAMB’s Convergence Speed. IEEE 28th International Conference on High Performance Computing, Data, and Analytics. 2022
2022
Cited alongside, same era.
Li, C.; Zhang, M.; He, Y. The Stability-Efficiency Dilemma: Investigating Sequence Length Warmup for Training GPT Models. Advances in Neural Information Processing Systems. 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Rajbhandari, S.; Li, C.; Yao, Z.; Zhang, M.; Aminabadi, R. Y.; Awan, A. A.; Rasley, J.; He, Y. DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale. International Conference on Machine Learning. 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Wu, X.; Yao, Z.; Zhang, M.; Li, C.; He, Y. Extreme Compression for Pre-trained Transformers Made Simple and Efficient. Advances in Neural Information Processing Systems. 2022
2022
Cited alongside, same era.
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Özçelik, R.; de Ruiter, S.; Criscuolo, E.; Grisoni, F. Chemical language modeling with structured state space sequence models. Nature Communications 2024
2024
Closest in time.
Li, J.; Zhang, O.; Sun, K.; Wang, Y.; Guan, X.; Bagni, D.; Haghighatlari, M.; Kearns, F. L.; Parks, C.; Amaro, R. E.; Head-Gordon, T. Mining for Potent Inhibitors through Artificial Intelligence and Physics: A Unified Methodology for Ligand Based and Structure Based Drug Design. Journal of Chemical Information and Modeling 2024
2024
Closest in time.
2024
Closest in time.
Templeton, A. et al. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Transformer Circuits Thread 2024
2024
Closest in time.
2024
Closest in time.
Mack, A.; Turner, A. Mechanistically Eliciting Latent Behaviors in Language Models. AI Alignment Forum 2024
2024
Closest in time.
Velez-Arce, A.; Huang, K.; Li, M.; Lin, X.; Gao, W.; Fu, T.; Kellis, M.; Pentelute, B. L.; Zitnik, M. TDC-2: Multimodal Foundation for Therapeutic Science. bioRxiv 2024
2024
Closest in time.
Dao, T. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. International Conference on Learning Representations (ICLR). 2024
2024
Closest in time.
2024
Closest in time.
Jacobs, S. A.; Tanaka, M.; Zhang, C.; Zhang, M.; Aminadabi, R. Y.; Song, S. L.; Rajbhandari, S.; He, Y. System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models. arXiv preprint 2024
2024
Closest in time.
2024
Closest in time.
Enamine Essential Fragment Library. https://enamine.net/compound-libraries/fragment-libraries/essential-library
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Bhattacharya, D.; Cassady, H.; Hickner, M.; Reinhart, W. Large Language Models as Molecular Design Engines. 2024; https://chemrxiv.org/engage/chemrxiv/article-details/664c98ea418a5379b0e07d31
2024
Closest in time.
2024
Closest in time.
Bagal, V.; Aggarwal, R.; Vinod, P. K.; Priyakumar, U. MolGPT: Molecular Generation Using a Transformer-Decoder Model. Journal of chemical information and modeling 2021
2076
Closest in time.