Fetching the paper…
Reading the bibliography…
Recent advancements in biological research leverage the integration of molecules, proteins, and natural language to enhance drug discovery.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Rapid and sensitive protein similarity searches
David J Lipman and William R Pearson. 1985 · 1985
Earlier work this paper cites.
Improved tools for biological sequence comparison
William R Pearson and David J Lipman. 1988 · 1988
Earlier work this paper cites.
Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules
David Weininger. 1988 · 1988
Earlier work this paper cites.
Smiles. 2. algorithm for generation of unique smiles notation
David Weininger, Arthur Weininger, and Joseph L Weininger. 1989 · 1989
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik. 1995 · 1995
Earlier work this paper cites.
Random decision forests
Tin Kam Ho. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Unsupervised data base clustering based on daylight’s fingerprint and tanimoto similarity: A fast and automated way to cluster small and large data sets
Darko Butina. 1999 · 1999
Earlier work this paper cites.
Machine learning in drug discovery: a review
Suresh Dara, Swetha Dhamercherla, Surender Singh Jadav, CH Madhu Babu, and Mohamed Jawed Ahsan. 2022 · 1999
Earlier work this paper cites.
Prediction of membrane protein types based on the hydrophobic index of amino acids
Zhi-Ping Feng and Chun-Ting Zhang. 2000 · 2000
Earlier work this paper cites.
Medical subject headings (mesh)
Carolyn E Lipscomb. 2000 · 2000
Earlier work this paper cites.
Recurrent neural networks
Larry R Medsker and LC Jain. 2001 · 2001
Earlier work this paper cites.
Reoptimization of mdl keys for use in drug discovery
Joseph L Durant, Burton A Leland, Douglas R Henry, and James G Nourse. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Development of human protein reference database as an initial platform for approaching systems biology in humans
Suraj Peri, J Daniel Navarro, Ramars Amanchy, Troels Z Kristiansen, Chandra Kiran Jonnalagadda, Vineeth Surendranath, Vidya Niranjan, Babylakshmi Muthusamy, TKB Gandhi, Mads Gronborg, et al. 2003 · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: an automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Uniprotkb/swiss-prot: the manually annotated section of the uniprot knowledgebase
Emmanuel Boutet, Damien Lieberherr, Michael Tognolli, Michel Schneider, and Amos Bairoch. 2007 · 2007
Earlier work this paper cites.
Bindingdb: a web-accessible database of experimentally determined protein–ligand binding affinities
Tiqing Liu, Yuhmei Lin, Xin Wen, Robert N Jorissen, and Michael K Gilson. 2007 · 2007
Earlier work this paper cites.
Uniref: comprehensive and non-redundant uniprot reference clusters
Baris E Suzek, Hongzhan Huang, Peter McGarvey, Raja Mazumder, and Cathy H Wu. 2007 · 2007
Earlier work this paper cites.
Using support vector machine combined with auto covariance to predict protein–protein interactions from protein sequences
Yanzhi Guo, Lezheng Yu, Zhining Wen, and Menglong Li. 2008 · 2008
Earlier work this paper cites.
Levenshtein distance: Information theory, computer science, string (computer science), string metric, damerau? levenshtein distance, spell checker, hamming distance
Frederic P Miller, Agnes F Vandome, and John McBrewster. 2009 · 2009
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Stephen E. Robertson and Hugo Zaragoza. 2009 · 2009
Earlier work this paper cites.
Chemberta: Large-scale self-supervised pretraining for molecular property prediction
Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. 2020 · 2010
Earlier work this paper cites.
Large-scale prediction of human protein- protein interactions from amino acid sequence based on latent topic features
Xiao-Yong Pan, Ya-Nan Zhang, and Hong-Bin Shen. 2010 · 2010
Earlier work this paper cites.
Pubmed: the bibliographic database
Kathi Canese and Sarah Weis. 2013 · 2013
Earlier work this paper cites.
propy: a tool to generate various modes of chou’s pseaac
Dong-Sheng Cao, Qing-Song Xu, and Yi-Zeng Liang. 2013 · 2013
Earlier work this paper cites.
Inchi-the worldwide chemical structure identifier standard
Stephen Heller, Alan McNaught, Stephen Stein, Dmitrii Tchekhovskoi, and Igor Pletnev. 2013 · 2013
Earlier work this paper cites.
Ncbi viral genomes resource
J Rodney Brister, Danso Ako-Adjei, Yiming Bao, and Olga Blinkova. 2015 · 2015
Earlier work this paper cites.
Improving compound–protein interaction prediction by building up highly credible negative samples
Hui Liu, Jianjiang Sun, Jihong Guan, Jie Zheng, and Shuigeng Zhou. 2015 · 2015
Cited alongside, same era.
An introduction to convolutional neural networks
Keiron O’Shea and Ryan Nash. 2015 · 2015
Cited alongside, same era.
Harnessing computational biology for exact linear b-cell epitope prediction: a novel amino acid composition-based feature descriptor
Vijayakumar Saravanan and Namasivayam Gautham. 2015 · 2015
Cited alongside, same era.
Get your atoms in order - an open-source implementation of a novel and robust molecular canonicalization algorithm
Nadine Schneider, Roger A. Sayle, and Gregory A. Landrum. 2015 · 2015
Cited alongside, same era.
ZINC 15 - ligand discovery for everyone
Teague Sterling and John J. Irwin. 2015 · 2015
Cited alongside, same era.
Text2mol: Cross-modal molecule retrieval with natural language queries
Carl Edwards, ChengXiang Zhai, and Heng Ji. 2021 · 2021
Later among the works it cites.
Prottrans: Toward understanding the language of life through self-supervised learning
Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, et al. 2021 · 2021
Later among the works it cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021 · 2021
Later among the works it cites.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2021 · 2021
Later among the works it cites.
MolTrans: Molecular interaction transformer for drug–target interaction prediction
Kexin Huang, Cao Xiao, Lucas Glass, and Jimeng Sun. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chebi in 2016: Improved services and an expanding collection of metabolites
Janna Hastings, Gareth Owen, Adriano Dekker, Marcus Ennis, Namrata Kale, Venkatesh Muthukrishnan, Steve Turner, Neil Swainston, Pedro Mendes, and Christoph Steinbeck. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Deeploc: prediction of protein subcellular localization using deep learning
José Juan Almagro Armenteros, Casper Kaae Sønderby, Søren Kaae Sønderby, Henrik Nielsen, and Ole Winther. 2017 · 2017
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Deepsol: a deep learning framework for sequence-based protein solubility prediction
Sameer Khurana, Reda Rawi, Khalid Kunji, Gwo-Yu Chuang, Halima Bensmail, and Raghvendra Mall. 2018 · 2018
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Rdkit: Open-source cheminformatics software
Greg Landrum. 2021 · 2021
Later among the works it cites.
GraphDTA: Predicting drug-target binding affinity with graph neural networks
Thin Nguyen, Hang Le, T. Quinn, Tri Minh Nguyen, Thuc Duy Le, and Svetha Venkatesh. 2021 · 2021
Later among the works it cites.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al. 2021 · 2021
Later among the works it cites.
Motif-based graph self-supervised learning for molecular property prediction
Zaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu, and Chee-Kong Lee. 2021 · 2021
Later among the works it cites.
Translation between molecules and natural language
Carl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji. 2022 · 2022
Later among the works it cites.
Geometry-enhanced molecular representation learning for property prediction
Xiaomin Fang, Lihang Liu, Jieqiong Lei, Donglong He, Shanzhuo Zhang, Jingbo Zhou, Fan Wang, Hua Wu, and Haifeng Wang. 2022 · 2022
Later among the works it cites.
Protgpt2 is a deep unsupervised language model for protein design
Noelia Ferruz, Steffen Schmidt, and Birte Höcker. 2022 · 2022
Later among the works it cites.
Selfies and the future of molecular string representations
Mario Krenn, Qianxiang Ai, Senja Barthel, Nessa Carson, Angelo Frei, Nathan C Frey, Pascal Friederich, Théophile Gaudin, Alberto Alexander Gayle, Kevin Maik Jablonka, et al. 2022 · 2022
Later among the works it cites.
Language models of protein sequences at the scale of evolution enable accurate structure prediction
Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Sal Candido, et al. 2022 · 2022
Later among the works it cites.
Pre-training molecular graph representation with 3d geometry
Shengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby, Hongyu Guo, and Jian Tang. 2022 · 2022
Later among the works it cites.
Biogpt: generative pre-trained transformer for biomedical text generation and mining
Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022 · 2022
Later among the works it cites.
A molecular multimodal foundation model associating molecule graphs with natural language
Bing Su, Dazhao Du, Zhao Yang, Yujie Zhou, Jiangmeng Li, Anyi Rao, Hao Sun, Zhiwu Lu, and Ji-Rong Wen. 2022 · 2022
Later among the works it cites.
Bern2: an advanced neural biomedical named entity recognition and normalization tool
Mujeen Sung, Minbyul Jeong, Yonghwa Choi, Donghyeon Kim, Jinhyuk Lee, and Jaewoo Kang. 2022 · 2022
Later among the works it cites.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022 · 2022
Later among the works it cites.
Molecular contrastive learning of representations via graph neural networks
Yuyang Wang, Jianren Wang, Zhonglin Cao, and Amir Barati Farimani. 2022 · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022 · 2022
Later among the works it cites.
Peer: a comprehensive and multi-task benchmark for protein sequence understanding
Minghao Xu, Zuobai Zhang, Jiarui Lu, Zhaocheng Zhu, Yangtian Zhang, Ma Chang, Runcheng Liu, and Jian Tang. 2022 · 2022
Later among the works it cites.
A deep-learning system bridging molecule structure and biomedical text with comprehension comparable to human professionals
Zheni Zeng, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2022 · 2022
Later among the works it cites.
Uniprot: the universal protein knowledgebase in 2023
2023 · 2023
Closest in time.
A big future for small molecules: targeting the undruggable
AstraZeneca. 2023 · 2023
Closest in time.
Interpretable bilinear attention network with domain adaptation improves drug–target prediction
Peizhen Bai, Filip Miljković, Bino John, and Haiping Lu. 2023 · 2023
Closest in time.
Pubchem 2023 update
Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al. 2023 · 2023
Closest in time.
Jiatong Li, Yunqing Liu, Wenqi Fan, Xiao-Yong Wei, Hui Liu, Jiliang Tang, and Qing Li. 2023 · 2023
Closest in time.
Molxpt: Wrapping molecules with text for generative pre-training
Zequn Liu, Wei Zhang, Yingce Xia, Lijun Wu, Shufang Xie, Tao Qin, Ming Zhang, and Tie-Yan Liu. 2023b · 2023
Closest in time.
Empowering ai drug discovery with explicit and implicit knowledge
Yizhen Luo, Kui Huang, Massimo Hong, Kai Yang, Jiahuan Zhang, Yushuai Wu, and Zaiqin Nie. 2023 · 2023
Closest in time.
Molgpt: Molecular generation using a transformer-decoder model
Viraj Bagal, Rishal Aggarwal, P. K. Vinod, and U. Deva Priyakumar. 2022 · 2076
Closest in time.