Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have emerged as a transformative power in enhancing natural language comprehension, representing a significant stride toward artificial general intelligence.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
SciBERT: A Pretrained Language Model for Scientific Text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 1903
Earlier work this paper cites.
ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. 2020 · 1904
Earlier work this paper cites.
Probing Biomedical Embeddings from Language Models
Qiao Jin, Bhuwan Dhingra, William W. Cohen, and Xinghua Lu. 2019a · 1904
Earlier work this paper cites.
PubMedQA: A Dataset for Biomedical Research Question Answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu. 2019b · 1909
Earlier work this paper cites.
S2ORC: The Semantic Scholar Open Research Corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Dan S. Weld. 2020 · 1911
Earlier work this paper cites.
Tractatus Logico-Philosophicus
Frank P Ramsey. 1923 · 1923
Earlier work this paper cites.
Nomenclature and symbolism for amino acids and peptides
H BIELKA GDR, N Sharon, and EW Australia. 1984 · 1984
Earlier work this paper cites.
SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules
David Weininger. 1988 · 1988
Earlier work this paper cites.
CATH–a hierarchic classification of protein domain structures
Christine A Orengo, Alex D Michie, Susan Jones, David T Jones, Mark B Swindells, and Janet M Thornton. 1997 · 1997
Earlier work this paper cites.
EDITtoTrEMBL: a distributed approach to high-quality automated protein sequence annotation
S Moller, Ulf Leser, Wolfgang Fleischmann, and Rolf Apweiler. 1999 · 1999
Earlier work this paper cites.
The mammalian gene collection
Robert L Strausberg, Elise A Feingold, Richard D Klausner, and Francis S Collins. 1999 · 1999
Earlier work this paper cites.
Gene ontology: tool for the unification of biology
Michael Ashburner, Catherine A Ball, Judith A Blake, David Botstein, Heather Butler, J Michael Cherry, Allan P Davis, Kara Dolinski, Selina S Dwight, Janan T Eppig, et al · 2000
Earlier work this paper cites.
The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000
Amos Bairoch and Rolf Apweiler. 2000 · 2000
Earlier work this paper cites.
Viral Genome DataBase: storing and analyzing genes and proteins from complete viral genomes
David Hiscock and Chris Upton. 2000 · 2000
Earlier work this paper cites.
SCOP: a structural classification of proteins database
Loredana Lo Conte, Bart Ailey, Tim JP Hubbard, Steven E Brenner, Alexey G Murzin, and Cyrus Chothia. 2000 · 2000
Earlier work this paper cites.
Nucleic acids: General properties
Garrett A Soukup. 2001 · 2001
Earlier work this paper cites.
A revision of Bloom’s taxonomy: An overview
David R Krathwohl. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
DIP, the Database of Interacting Proteins: a research tool for studying cellular networks of protein interactions
Ioannis Xenarios, Lukasz Salwinski, Xiaoqun Joyce Duan, Patrick Higney, Sul-Min Kim, and David Eisenberg. 2002 · 2002
Earlier work this paper cites.
The UCSC genome browser database
Donna Karolchik, Robert Baertsch, Mark Diekhans, Terrence S Furey, Angie Hinrichs, YT Lu, Krishna M Roskin, Matt Schwartz, Charles W Sugnet, Daryl J Thomas, et al · 2003
Earlier work this paper cites.
The ENCODE (ENCyclopedia of DNA elements) project
EA Feingold, PJ Good, MS Guyer, S Kamholz, L Liefer, K Wetterstrand, FS Collins, TR Gingeras, D Kampa, EA Sekinger, et al · 2004
Earlier work this paper cites.
MedDialog: Two Large-scale Medical Dialogue Datasets
Xuehai He, Shu Chen, Zeqian Ju, Xiangyu Dong, Hongchao Fang, Sicheng Wang, Yue Yang, Jiaqi Zeng, Ruisi Zhang, Ruoyu Zhang, Meng Zhou, Penghui Zhu, and Pengtao Xie. 2020 · 2004
Earlier work this paper cites.
The Wave in the Mind: Talks and Essays on the Writer, the Reader, and the Imagination
Ursula K Le Guin. 2004 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
ProGen: Language Modeling for Protein Generation
Ali Madani, Bryan McCann, Nikhil Naik, Nitish Shirish Keskar, Namrata Anand, Raphael R. Eguchi, Po-Ssu Huang, and Richard Socher. 2020 · 2004
Earlier work this paper cites.
PIRSF: family classification system at the Protein Information Resource
Cathy H Wu, Anastasia Nikolskaya, Hongzhan Huang, Lai-Su L Yeh, Darren A Natale, Cholanayakanahalli R Vinayaka, Zhang-Zhi Hu, Raja Mazumder, Sandeep Kumar, Panagiotis Kourtesis, et al · 2004
Earlier work this paper cites.
An Empirical Study of Multi-Task Learning on BERT for Biomedical Text Mining
Yifan Peng, Qingyu Chen, and Zhiyong Lu. 2020 · 2005
Earlier work this paper cites.
The PDBbind database: methodologies and updates
Renxiao Wang, Xueliang Fang, Yipin Lu, Chao-Yie Yang, and Shaomeng Wang. 2005 · 2005
Earlier work this paper cites.
Pfam: clans, web tools and services
Robert D Finn, Jaina Mistry, Benjamin Schuster-Böckler, Sam Griffiths-Jones, Volker Hollich, Timo Lassmann, Simon Moxon, Mhairi Marshall, Ajay Khanna, Richard Durbin, et al · 2006
Earlier work this paper cites.
UniProtKB/Swiss-Prot: the manually annotated section of the UniProt KnowledgeBase
Emmanuel Boutet, Damien Lieberherr, Michael Tognolli, Michel Schneider, and Amos Bairoch. 2007 · 2007
Earlier work this paper cites.
UniRef: comprehensive and non-redundant UniProt reference clusters
Baris E Suzek, Hongzhan Huang, Peter McGarvey, Raja Mazumder, and Cathy H Wu. 2007 · 2007
Earlier work this paper cites.
Database resources of the national center for biotechnology information
David L Wheeler, Tanya Barrett, Dennis A Benson, Stephen H Bryant, Kathi Canese, Vyacheslav Chetvernin, Deanna M Church, Michael DiCuccio, Ron Edgar, Scott Federhen, et al · 2007
Earlier work this paper cites.
Scoring function for automated assessment of protein structure template quality
Yang Zhang and Jeffrey Skolnick. 2007 · 2007
Earlier work this paper cites.
Protein structure–structure alignment with discrete Fréchet distance
Minghui Jiang, Ying Xu, and Binhai Zhu. 2008 · 2008
Earlier work this paper cites.
ExplorEnz: the primary source of the IUBMB enzyme list
Andrew G. McDonald, Sinéad Boyce, and Keith F. Tipton. 2008 · 2008
Earlier work this paper cites.
970 million druglike small molecules for virtual screening in the chemical universe database GDB-13
Lorenz C Blum and Jean-Louis Reymond. 2009 · 2009
Earlier work this paper cites.
The myth of language universals: Language diversity and its importance for cognitive science
Nicholas Evans and Stephen C Levinson. 2009 · 2009
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2009
Earlier work this paper cites.
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2020 · 2009
Earlier work this paper cites.
PubChem: a public information system for analyzing bioactivities of small molecules
Yanli Wang, Jewen Xiao, Tugba O Suzek, Jian Zhang, Jiyao Wang, and Stephen H Bryant. 2009 · 2009
Earlier work this paper cites.
Bloom’s taxonomy
Mary Forehand. 2010 · 2010
Earlier work this paper cites.
The RCSB Protein Data Bank: redesigned web site and web services
Peter W Rose, Bojan Beran, Chunxiao Bi, Wolfgang F Bluhm, Dimitris Dimitropoulos, David S Goodsell, Andreas Prlić, Martha Quesada, Gregory B Quinn, John D Westbrook, et al · 2010
Earlier work this paper cites.
BioMegatron: Larger Biomedical Domain Language Model
Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, and Raghav Mani. 2020 · 2010
Earlier work this paper cites.
An integrated encyclopedia of DNA elements in the human genome
ENCODE Project Consortium Overall coordination (data analysis coordination) Dunham Ian 2 Kundaje Anshul 3 81 82 82, Writing group Bernstein Bradley E. 7 34 Birney Ewan Dunham Ian Green Eric D. 35 Gunter Chris 15 Snyder Michael 13, et al · 2012
Earlier work this paper cites.
ChEMBL: a large-scale bioactivity database for drug discovery
Anna Gaulton, Louisa J Bellis, A Patricia Bento, Jon Chambers, Mark Davies, Anne Hersey, Yvonne Light, Shaun McGlinchey, David Michalovich, Bissan Al-Lazikani, et al · 2012
Earlier work this paper cites.
ZINC: a free tool to discover chemistry for biology
John J Irwin, Teague Sterling, Michael M Mysinger, Erin S Bolstad, and Ryan G Coleman. 2012 · 2012
Earlier work this paper cites.
The Arabidopsis Information Resource (TAIR): improved gene annotation and new tools
Philippe Lamesch, Tanya Z Berardini, Donghui Li, David Swarbreck, Christopher Wilks, Rajkumar Sasidharan, Robert Muller, Kate Dreher, Debbie L Alexander, Margarita Garcia-Hernandez, et al · 2012
Earlier work this paper cites.
Extraction of chemical structures and reactions from the literature
Daniel Mark Lowe. 2012 · 2012
Earlier work this paper cites.
Directory of useful decoys, enhanced (DUD-E): better ligands and decoys for better benchmarking
Michael M Mysinger, Michael Carchia, John J Irwin, and Brian K Shoichet. 2012 · 2012
Earlier work this paper cites.
Enumeration of 166 billion organic small molecules in the chemical universe database GDB-17
Lars Ruddigkeit, Ruud Van Deursen, Lorenz C Blum, and Jean-Louis Reymond. 2012 · 2012
Earlier work this paper cites.
HIPPIE: Integrating protein interaction networks with experiment based quality scores
Martin H Schaefer, Jean-Fred Fontaine, Arunachalam Vinayagam, Pablo Porras, Erich E Wanker, and Miguel A Andrade-Navarro. 2012 · 2012
Earlier work this paper cites.
New and continuing developments at PROSITE
Christian JA Sigrist, Edouard De Castro, Lorenzo Cerutti, Béatrice A Cuche, Nicolas Hulo, Alan Bridge, Lydie Bougueleret, and Ioannis Xenarios. 2012 · 2012
Earlier work this paper cites.
PubChem’s BioAssay database
Yanli Wang, Jewen Xiao, Tugba O Suzek, Jian Zhang, Jiyao Wang, Zhigang Zhou, Lianyi Han, Karen Karapetyan, Svetlana Dracheva, Benjamin A Shoemaker, et al · 2012
Earlier work this paper cites.
BioLiP: a semi-manually curated database for biologically relevant ligand–protein interactions
Jianyi Yang, Ambrish Roy, and Yang Zhang. 2012 · 2012
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. 2013 · 2013
Earlier work this paper cites.
PubMed: the bibliographic database
Kathi Canese and Sarah Weis. 2013 · 2013
Earlier work this paper cites.
EPD and EPDnew, high-quality promoter resources in the next-generation sequencing era
René Dreos, Giovanna Ambrosini, Rouayda Cavin Périer, and Philipp Bucher. 2013 · 2013
Earlier work this paper cites.
Nomenclature of organic chemistry: IUPAC recommendations and preferred names 2013
Henri A Favre and Warren H Powell. 2013 · 2013
Earlier work this paper cites.
InChI- the worldwide chemical structure identifier standard
Stephen Heller, Alan McNaught, Stephen Stein, Dmitrii Tchekhovskoi, and Igor Pletnev. 2013 · 2013
Earlier work this paper cites.
lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests
Valerio Mariani, Marco Biasini, Alessandro Barbato, and Torsten Schwede. 2013 · 2013
Earlier work this paper cites.
A global reference for human genetic variation
1000 Genomes Project Consortium et al · 2015
Earlier work this paper cites.
ZINC 15–ligand discovery for everyone
Teague Sterling and John J Irwin. 2015 · 2015
Earlier work this paper cites.
UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches
Baris E Suzek, Yuqi Wang, Hongzhan Huang, Peter B McGarvey, Cathy H Wu, and UniProt Consortium. 2015 · 2015
Earlier work this paper cites.
Predicting effects of noncoding variants with deep learning–based sequence model
Jian Zhou and Olga G Troyanskaya. 2015 · 2015
Earlier work this paper cites.
UniProtKB/Swiss-Prot, the manually annotated section of the UniProt KnowledgeBase: how to use the entry view
Emmanuel Boutet, Damien Lieberherr, Michael Tognolli, Michel Schneider, Parit Bansal, Alan J Bridge, Sylvain Poux, Lydie Bougueleret, and Ioannis Xenarios. 2016 · 2016
Earlier work this paper cites.
The gene expression omnibus database
Emily Clough and Tanya Barrett. 2016 · 2016
Earlier work this paper cites.
Preparing a collection of radiology examinations for distribution and retrieval
Dina Demner-Fushman, Marc D Kohli, Marc B Rosenman, Sonya E Shooshan, Laritza Rodriguez, Sameer Antani, George R Thoma, and Clement J McDonald. 2016 · 2016
Earlier work this paper cites.
BindingDB in 2015: a public database for medicinal chemistry, computational chemistry and systems pharmacology
Michael K Gilson, Tiqing Liu, Michael Baitaluk, George Nicola, Linda Hwang, and Jenny Chong. 2016 · 2016
Earlier work this paper cites.
ChEBI in 2016: Improved services and an expanding collection of metabolites
Janna Hastings, Gareth Owen, Adriano Dekker, Marcus Ennis, Namrata Kale, Venkatesh Muthukrishnan, Steve Turner, Neil Swainston, Pedro Mendes, and Christoph Steinbeck. 2016 · 2016
Earlier work this paper cites.
MIMIC-III, a freely accessible critical care database
Alistair E. W. Johnson, Tom J. Pollard, Lu Shen, Li wei H. Lehman, Mengling Feng, Mohammad Mahdi Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G. Mark. 2016 · 2016
Earlier work this paper cites.
PubChem substance and compound databases
Sunghwan Kim, Paul A Thiessen, Evan E Bolton, Jie Chen, Gang Fu, Asta Gindulyte, Lianyi Han, Jane He, Siqian He, Benjamin A Shoemaker, et al · 2016
Earlier work this paper cites.
ChemProt-3.0: a global chemical biology diseases mapping
Jens Kringelum, Sonny Kim Kjaerulff, Søren Brunak, Ole Lund, Tudor I Oprea, and Olivier Taboureau. 2016 · 2016
Earlier work this paper cites.
BioCreative V CDR task corpus: a resource for chemical disease relation extraction
Jiao Li, Yueping Sun, Robin J Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J Mattingly, Thomas C Wiegers, and Zhiyong Lu. 2016 · 2016
Earlier work this paper cites.
What’s what: The (nearly) definitive guide to reaction role assignment
Nadine Schneider, Nikolaus Stiefl, and Gregory A Landrum. 2016 · 2016
Earlier work this paper cites.
The FAIR Guiding Principles for scientific data management and stewardship
Mark D Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Appleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan-Willem Boiten, Luiz Bonino da Silva Santos, Philip E Bourne, et al · 2016
Earlier work this paper cites.
Convolutional neural network architectures for predicting DNA–protein binding
Haoyang Zeng, Matthew D Edwards, Ge Liu, and David K Gifford. 2016 · 2016
Earlier work this paper cites.
ChemGAN challenge for drug discovery: can AI reproduce natural chemical diversity?
Mostapha Benhenda. 2017 · 2017
Earlier work this paper cites.
Mouse Genome Informatics (MGI): resources for mining mouse genetic, genomic, and biological data in support of primary and translational research
Janan T Eppig, Cynthia L Smith, Judith A Blake, Martin Ringwald, James A Kadin, Joel E Richardson, and Carol J Bult. 2017 · 2017
Earlier work this paper cites.
Improvements and impacts of GRCh38 human reference on high throughput sequencing data analysis
Yan Guo, Yulin Dai, Hui Yu, Shilin Zhao, David C Samuels, and Yu Shyr. 2017 · 2017
Earlier work this paper cites.
Crowdsourcing Multiple Choice Science Questions
Matt Gardner Johannes Welbl, Nelson F. Liu. 2017 · 2017
Earlier work this paper cites.
The human cell atlas
Aviv Regev, Sarah A Teichmann, Eric S Lander, Ido Amit, Christophe Benoist, Ewan Birney, Bernd Bodenmiller, Peter Campbell, Piero Carninci, Menna Clatworthy, et al · 2017
Earlier work this paper cites.
ExCAPE-DB: an integrated large scale dataset facilitating Big Data analysis in chemogenomics
Jiangming Sun, Nina Jeliazkova, Vladimir Chupakhin, Jose-Felipe Golib-Dzib, Ola Engkvist, Lars Carlsson, Jörg Wegner, Hugo Ceulemans, Ivan Georgiev, Vedrin Jeliazkov, et al · 2017
Earlier work this paper cites.
Attention is All you Need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F Liu, and Matt Gardner. 2017 · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Massive mining of publicly available RNA-seq data from human and mouse
Alexander Lachmann, Denis Torre, Alexandra B Keenan, Kathleen M Jagodnik, Hoyjin J Lee, Lily Wang, Moshe C Silverstein, and Avi Ma’ayan. 2018 · 2018
Earlier work this paper cites.
The eICU Collaborative Research Database, a freely available multi-center database for critical care research
Tom J Pollard, Alistair EW Johnson, Jesse D Raffa, Leo A Celi, Roger G Mark, and Omar Badawi. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Sequence-based prediction of variants’ effects
Nicole Rusk. 2018 · 2018
Earlier work this paper cites.
Modeling relational data with graph convolutional networks. In The Semantic Web: 15th International Conference, ESWC 2018, Heraklion, Crete, Greece, June 3–7, 2018, Proceedings 15 . Springer, 593–607
Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne Van Den Berg, Ivan Titov, and Max Welling. 2018 · 2018
Earlier work this paper cites.
Clustering huge protein sequence sets in linear time
Martin Steinegger and Johannes Söding. 2018 · 2018
Earlier work this paper cites.
DrugBank 5.0: a major update to the DrugBank database for 2018
David S Wishart, Yannick D Feunang, An C Guo, Elvis J Lo, Ana Marcu, Jason R Grant, Tanvir Sajed, Daniel Johnson, Carin Li, Zinat Sayeeda, et al · 2018
Earlier work this paper cites.
MoleculeNet: a benchmark for molecular machine learning
Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. 2018 · 2018
Earlier work this paper cites.
Protein Data Bank: the single global archive for 3D macromolecular structure data
wwPDB consortium. 2018 · 2018
Earlier work this paper cites.
How powerful are graph neural networks?
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. 2018 · 2018
Earlier work this paper cites.
Multi-scale attentive interaction networks for chinese medical question answer selection
Sheng Zhang, Xin Zhang, Hui Wang, Lixiang Guo, and Shanshan Liu. 2018 · 2018
Earlier work this paper cites.
A question-entailment approach to question answering
Asma Ben Abacha and Dina Demner-Fushman. 2019 · 2019
Earlier work this paper cites.
GuacaMol: benchmarking models for de novo molecular design
Nathan Brown, Marco Fiscato, Marwin HS Segler, and Alain C Vaucher. 2019 · 2019
Earlier work this paper cites.
CAGI 5 splicing challenge: improved exon skipping and intron retention predictions with MMSplice
Jun Cheng, Muhammed Hasan Çelik, Thi Yen Duong Nguyen, Žiga Avsec, and Julien Gagneur. 2019 · 2019
Earlier work this paper cites.
UniProt: a worldwide hub of protein knowledge
UniProt Consortium. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
PanglaoDB: a web server for exploration of mouse and human single-cell RNA sequencing data
Oscar Franzén, Li-Ming Gan, and Johan LM Björkegren. 2019 · 2019
Earlier work this paper cites.
Smiles transformer: Pre-trained molecular fingerprint for low data drug discovery
Shion Honda, Shoi Shi, and Hiroki R Ueda. 2019 · 2019
Earlier work this paper cites.
MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports
Alistair E. W. Johnson, Tom J. Pollard, Seth J. Berkowitz, Nathaniel R. Greenbaum, Matthew P. Lungren, Chih ying Deng, Roger G. Mark, and Steven Horng. 2019 · 2019
Earlier work this paper cites.
A transformer model for retrosynthesis. In International Conference on Artificial Neural Networks . Springer, 817–830
Pavel Karpov, Guillaume Godin, and Igor V Tetko. 2019 · 2019
Earlier work this paper cites.
CAGI5: Objective performance assessments of predictions based on the Evolutionary Action equation
Panagiotis Katsonis and Olivier Lichtarge. 2019 · 2019
Earlier work this paper cites.
PubChem 2019 update: improved access to chemical data
Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
The BioGRID interaction database: 2019 update
Rose Oughtred, Chris Stark, Bobby-Joe Breitkreutz, Jennifer Rust, Lorrie Boucher, Christie Chang, Nadine Kolas, Lara O’Donnell, Genie Leung, Rochelle McAdam, et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Evaluating Protein Transfer Learning with TAPE. In Advances in Neural Information Processing Systems
Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Xi Chen, John Canny, Pieter Abbeel, and Yun S Song. 2019 · 2019
Earlier work this paper cites.
Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences
Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C. Lawrence Zitnick, Jerry Ma, and Rob Fergus. 2019 · 2019
Earlier work this paper cites.
Molecular transformer: a model for uncertainty-calibrated chemical reaction prediction
Philippe Schwaller, Teodoro Laino, Théophile Gaudin, Peter Bolgar, Christopher A Hunter, Costas Bekas, and Alpha A Lee. 2019 · 2019
Earlier work this paper cites.
Incorporating domain knowledge into medical NLI using knowledge graphs
Soumya Sharma, Bishal Santra, Abhik Jana, TYSS Santosh, Niloy Ganguly, and Pawan Goyal. 2019 · 2019
Earlier work this paper cites.
Protein-level assembly increases protein sequence recovery from metagenomic samples manyfold
Martin Steinegger, Milot Mirdita, and Johannes Söding. 2019 · 2019
Earlier work this paper cites.
Smiles-bert: large scale unsupervised pre-training for molecular property prediction. In Proceedings of the 10th ACM international conference on bioinformatics, computational biology and health informatics . 429–436
Sheng Wang, Yuzhi Guo, Yuhong Wang, Hongmao Sun, and Junzhou Huang. 2019 · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Earlier work this paper cites.
Predicting retrosynthetic reactions using self-corrected transformer neural networks
Shuangjia Zheng, Jiahua Rao, Zhongyue Zhang, Jun Xu, and Yuedong Yang. 2019 · 2019
Earlier work this paper cites.
ChemBERTa: large-scale self-supervised pretraining for molecular property prediction
Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. 2020 · 2020
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020 · 2020
Earlier work this paper cites.
Molecular representation learning with language models and domain-relevant auxiliary tasks
Benedek Fabian, Thomas Edlich, Héléna Gaspar, Marwin Segler, Joshua Meyers, Marco Fiscato, and Mohamed Ahmed. 2020 · 2020
Earlier work this paper cites.
Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks. In Proceedings of ACL
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Open graph benchmark: Datasets for machine learning on graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. 2020 · 2020
Earlier work this paper cites.
ZINC20—a free ultralarge-scale chemical database for ligand discovery
John J Irwin, Khanh G Tang, Jennifer Young, Chinzorig Dandarchuluun, Benjamin R Wong, Munkhzul Khurelbaatar, Yurii S Moroz, John Mayfield, and Roger A Sayle. 2020 · 2020
Earlier work this paper cites.
The reactome pathway knowledgebase
Bijay Jassal, Lisa Matthews, Guilherme Viteri, Chuqiao Gong, Pascual Lorente, Antonio Fabregat, Konstantinos Sidiropoulos, Justin Cook, Marc Gillespie, Robin Haw, et al · 2020
Earlier work this paper cites.
Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation
Mario Krenn, Florian Häse, AkshatKumar Nigam, Pascal Friederich, and Alan Aspuru-Guzik. 2020 · 2020
Earlier work this paper cites.
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020 · 2020
Earlier work this paper cites.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 . 7871–7880
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Earlier work this paper cites.
Molecule attention transformer
Łukasz Maziarka, Tomasz Danel, Sławomir Mucha, Krzysztof Rataj, Jacek Tabor, and Stanisław Jastrzębski. 2020 · 2020
Cited alongside, same era.
Molecular sets (MOSES): a benchmarking platform for molecular generation models
Daniil Polykovskiy, Alexander Zhebrak, Benjamin Sanchez-Lengeling, Sergey Golovanov, Oktai Tatanov, Stanislav Belyaev, Rauf Kurbanov, Aleksey Artamonov, Vladimir Aladinskiy, Mark Veselov, et al · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Self-supervised graph transformer on large-scale molecular data
Yu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie, Ying Wei, Wenbing Huang, and Junzhou Huang. 2020 · 2020
Cited alongside, same era.
State-of-the-art augmented NLP transformer models for direct and single-step retrosynthesis
Igor V Tetko, Pavel Karpov, Ruud Van Deursen, and Guillaume Godin. 2020 · 2020
EpiGePT: a Pretrained Transformer model for epigenomics
Zijing Gao, Qiao Liu, Wanwen Zeng, Wing H Wong, and Rui Jiang. 2023a · 2023
Later among the works it cites.
Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge Evaluation
Zhouhong Gu, Xiaoxuan Zhu, Haoning Ye, Lin Zhang, Jianchen Wang, Sihang Jiang, Zhuozhi Xiong, Zihan Li, Qianyu He, Rui Xu, Wenhao Huang, Zili Wang, Shusen Wang, Weiguo Zheng, Hongwei Feng, and Yanghua Xiao. 2023 · 2023
Later among the works it cites.
Graph-based molecular representation learning
Zhichun Guo, Kehan Guo, Bozhao Nan, Yijun Tian, Roshni G Iyer, Yihong Ma, Olaf Wiest, Xiangliang Zhang, Wei Wang, Chuxu Zhang, et al · 2023
Later among the works it cites.
ProstT5: Bilingual Language Model for Protein Sequence and Structure
Michael Heinzinger, Konstantin Weissenow, Joaquin Gomez Sanchez, Adrian Henkel, Martin Steinegger, and Burkhard Rost. 2023 · 2023
Later among the works it cites.
The Diminishing Returns of Masked Language Models to Science
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
PubMed 2.0
Jacob White. 2020 · 2020
Cited alongside, same era.
X-MOL: large-scale pre-training for molecular understanding and diverse molecular analysis
Dongyu Xue, Han Zhang, Dongling Xiao, Yukang Gong, Guohui Chuai, Yu Sun, Hao Tian, Hua Wu, Yukun Li, and Qi Liu. 2020 · 2020
Cited alongside, same era.
Generative pre-training from molecules
Sanjar Adilov. 2021 · 2021
Cited alongside, same era.
BioM-Transformers: Building Large Biomedical Language Models with BERT, ALBERT and ELECTRA. In Proceedings of the 20th Workshop on Biomedical Language Processing . 221–227
Sultan Alrowili and Vijay Shanker. 2021 · 2021
Cited alongside, same era.
Effective gene expression prediction from sequence by integrating long-range interactions
Žiga Avsec, Vikram Agarwal, Daniel Visentin, Joseph R Ledsam, Agnieszka Grabska-Barwinska, Kyle R Taylor, Yannis Assael, John Jumper, Pushmeet Kohli, and David R Kelley. 2021 · 2021
Cited alongside, same era.
Learning the protein language: Evolution, structure, and function
Tristan Bepler and Bonnie Berger. 2021 · 2021
Cited alongside, same era.
Fold2seq: A joint sequence (1d)-fold (3d) embedding-based generative model for protein design. In International Conference on Machine Learning . PMLR, 1261–1271
Yue Cao, Payel Das, Vijil Chenthamarakshan, Pin-Yu Chen, Igor Melnyk, and Yang Shen. 2021 · 2021
Cited alongside, same era.
Zhi Hong, Aswathy Ajith, Gregory Pauloski, Eamon Duede, Kyle Chard, and Ian Foster. 2023 · 2023
Later among the works it cites.
A systematic benchmark of machine learning methods for protein–RNA interaction prediction
Marc Horlacher, Giulia Cantini, Julian Hesse, Patrick Schinke, Nicolas Goedert, Shubhankar Londhe, Lambert Moyon, and Annalisa Marsico. 2023 · 2023
Later among the works it cites.
Ogb-lsc: A large-scale challenge for machine learning on graphs
Weihua Hu, Matthias Fey, Hongyu Ren, Maho Nakata, Yuxiao Dong, and Jure Leskovec. 2023 · 2023
Later among the works it cites.
C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models. In Advances in Neural Information Processing Systems
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. 2023 · 2023
Later among the works it cites.
Language models in molecular discovery
Nikita Janakarajan, Tim Erdmann, Sarath Swaminathan, Teodoro Laino, and Jannis Born. 2023 · 2023
Later among the works it cites.
MIMIC-IV, a freely accessible electronic health record dataset
Alistair E. W. Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J. Pollard, Benjamin Moody, Brian Gow, Li wei H. Lehman, Leo Anthony Celi, and Roger G. Mark. 2023a · 2023
Later among the works it cites.
PubChem 2023 update
Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al · 2023
Later among the works it cites.
Pre-training Sequence, Structure, and Surface Features for Comprehensive Protein Representation Learning. In The Twelfth International Conference on Learning Representations
Youhan Lee, Hasun Yu, Jaemyung Lee, and Jaehoon Kim. 2023 · 2023
Later among the works it cites.
Cell2sentence: Teaching large language models the language of biology
Daniel Levine, Sacha Lévy, Syed Asad Rizvi, Nazreen Pallikkavaliyaveetil, Xingyu Chen, David Zhang, Sina Ghadermarzi, Ruiming Wu, Zihe Zheng, Ivan Vrkic, et al · 2023
Later among the works it cites.
BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023c · 2023
Later among the works it cites.
iEnhancer-ELM: improve enhancer identification by extracting position-related multiscale contextual information based on enhancer language models
Jiahao Li, Zhourun Wu, Wenhao Lin, Jiawei Luo, Jun Zhang, Qingcai Chen, and Junjie Chen. 2023g · 2023
Later among the works it cites.
Druggpt: A gpt-based strategy for designing potential ligands targeting specific proteins
Yuesen Li, Chengyi Gao, Xin Song, Xiangyu Wang, Yungang Xu, and Suxia Han. 2023a · 2023
Later among the works it cites.
PLPMpro: Enhancing promoter sequence prediction with prompt-learning based pre-trained language model
Zhongshen Li, Junru Jin, Wentao Long, and Leyi Wei. 2023b · 2023
Later among the works it cites.
DrugChat: towards enabling ChatGPT-like capabilities on drug molecule graphs
Youwei Liang, Ruiyi Zhang, Li Zhang, and Pengtao Xie. 2023 · 2023
Later among the works it cites.
MolRoPE-BERT: An enhanced molecular representation with Rotary Position Embedding for molecular property prediction
Yunwu Liu, Ruisheng Zhang, Tongfeng Li, Jing Jiang, Jun Ma, and Ping Wang. 2023e · 2023
Later among the works it cites.
Improving language model of human genome for DNA–protein binding prediction based on task-specific pre-training
Hanyu Luo, Wenyu Shan, Cheng Chen, Pingjian Ding, and Lingyun Luo. 2023a · 2023
Later among the works it cites.
Biomedgpt: Open multimodal generative pre-trained transformer for biomedicine
Yizhen Luo, Jiahuan Zhang, Siqi Fan, Kai Yang, Yushuai Wu, Mu Qiao, and Zaiqing Nie. 2023c · 2023
Later among the works it cites.
Rachel K. Luu and Markus J. Buehler. 2023 · 2023
Later among the works it cites.
Retrieved Sequence Augmentation for Protein Representation Learning
Chang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin, Qintong Li, Lijun Wu, Zhihong Deng, Yang Lu, Qi Liu, and Lingpeng Kong. 2023 · 2023
Later among the works it cites.
Aditya Malusare, Harish Kothandaraman, Dipesh Tamboli, Nadia A. Lanman, and Vaneet Aggarwal. 2023 · 2023
Later among the works it cites.
Ensembl 2023
Fergal J Martin, M Ridwan Amode, Alisha Aneja, Olanrewaju Austine-Orimoloye, Andrey G Azov, If Barnes, Arne Becker, Ruth Bennett, Andrew Berry, Jyothish Bhai, et al · 2023
Later among the works it cites.
Molecule generation using transformers and policy gradient reinforcement learning
Eyal Mazuz, Guy Shtar, Bracha Shapira, and Lior Rokach. 2023 · 2023
Later among the works it cites.
Orca 2: Teaching Small Language Models How to Reason
Arindam Mitra, Luciano Del Corro, Shweti Mahajan, Andres Codas, Clarisse Simoes, Sahaj Agrawal, Xuxi Chen, Anastasia Razdaibiedina, Erik Jones, Kriti Aggarwal, Hamid Palangi, Guoqing Zheng, Corby Rosset, Hamed Khanpour, and Ahmed Awadallah. 2023 · 2023
Later among the works it cites.
Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution
Eric Nguyen, Michael Poli, Marjan Faizi, Armin Thomas, Callum Birch-Sykes, Michael Wornow, Aman Patel, Clayton Rabideau, Stefano Massaroli, Yoshua Bengio, et al · 2023
Later among the works it cites.
ProGen2: exploring the boundaries of protein language models
Erik Nijkamp, Jeffrey A Ruffolo, Eli N Weinstein, Nikhil Naik, and Ali Madani. 2023 · 2023
Later among the works it cites.
ProteinGym: Large-Scale Benchmarks for Protein Design and Fitness Prediction
Pascal Notin, Aaron W. Kollasch, Daniel Ritter, Lood van Niekerk, Steffanie Paul, Hansen Spinner, Nathan Rollins, Ada Shaw, Ruben Weitzman, Jonathan Frazer, Mafalda Dias, Dinko Franceschi, Rose Orenbuch, Yarin Gal, and Debora S. Marks. 2023a · 2023
Later among the works it cites.
ProteinNPT: Improving Protein Property Prediction and Design with Non-Parametric Transformers
Pascal Notin, Ruben Weitzman, Debora S. Marks, and Yarin Gal. 2023b · 2023
Later among the works it cites.
Introducing the bacterial and viral bioinformatics resource center (BV-BRC): a resource combining PATRIC, IRD and ViPR
Robert D Olson, Rida Assaf, Thomas Brettin, Neal Conrad, Clark Cucinell, James J Davis, Donald M Dempsey, Allan Dickerman, Emily M Dietrich, Ronald W Kenyon, et al · 2023
Later among the works it cites.
GPT-4 Technical Report
OpenAI. 2023 · 2023
Later among the works it cites.
InterPro in 2022
Typhaine Paysan-Lafosse, Matthias Blum, Sara Chuguransky, Tiago Grego, Beatriz Lázaro Pinto, Gustavo A Salazar, Maxwell L Bileschi, Peer Bork, Alan Bridge, Lucy Colwell, et al · 2023
Later among the works it cites.
Qizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu, Kaiyuan Gao, Lijun Wu, Yingce Xia, and Rui Yan. 2023 · 2023
Later among the works it cites.
A study of generative large language model for medical research and healthcare
Cheng Peng, Xi Yang, Aokun Chen, Kaleb E. Smith, Nima PourNejatian, Anthony B. Costa, Cheryl Martin, Mona G. Flores, Ying Zhang, Tanja Magoc, Gloria Lipori, Duane A. Mitchell, Naykky S. Ospina, Mustafa M. Ahmed, William R. Hogan, Elizabeth A. Shenkman, Yi Guo, Jiang Bian, and Yonghui Wu. 2023 · 2023
Later among the works it cites.
Efficient and accurate sequence generation with small-scale protein language models
Yaiza Serrano, Sergi Roda, Victor Guallar, and Alexis Molina. 2023 · 2023
Later among the works it cites.
Generative power of a protein language model trained on multiple sequence alignments
Damiano Sgarbossa, Umberto Lupo, and Anne-Florence Bitbol. 2023 · 2023
Later among the works it cites.
A general-purpose material property data extraction pipeline from large polymer corpora using natural language processing
Pranav Shetty, Arunkumar Chitteth Rajan, Chris Kuenneth, Sonakshi Gupta, Lakshmi Prerana Panchumarti, Lauren Holm, Chao Zhang, and Rampi Ramprasad. 2023 · 2023
Later among the works it cites.
CPLLM: Clinical Prediction with Large Language Models
Ofir Ben Shoham and Nadav Rappoport. 2023 · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2023
Later among the works it cites.
ProteinRL: Reinforcement learning with generative protein language models for property-directed sequence design. In NeurIPS 2023 Generative AI and Biology (GenBio) Workshop
Matt Sternke and Joel Karpiak. 2023 · 2023
Later among the works it cites.
Saprot: Protein language modeling with structure-aware vocabulary
Jin Su, Chenchen Han, Yuyang Zhou, Junjie Shan, Xibin Zhou, and Fajie Yuan. 2023 · 2023
Later among the works it cites.
SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research
Liangtai Sun, Yang Han, Zihan Zhao, Da Ma, Zhennan Shen, Baocai Chen, Lu Chen, and Kai Yu. 2023 · 2023
Later among the works it cites.
The STRING database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest
Damian Szklarczyk, Rebecca Kirsch, Mikaela Koutrouli, Katerina Nastou, Farrokh Mehryary, Radja Hachilif, Annika L Gable, Tao Fang, Nadezhda T Doncheva, Sampo Pyysalo, et al · 2023
Later among the works it cites.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team. 2023 · 2023
Later among the works it cites.
Unbiasing retrosynthesis language models with disconnection prompts
Amol Thakkar, Alain C Vaucher, Andrea Byekwaso, Philippe Schwaller, Alessandra Toniato, and Teodoro Laino. 2023 · 2023
Later among the works it cites.
Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding
Augustin Toma, Patrick R. Lawler, Jimmy Ba, Rahul G. Krishnan, Barry B. Rubin, and Bo Wang. 2023 · 2023
Later among the works it cites.
Enhancing diversity in language based models for single-step retrosynthesis
Alessandra Toniato, Alain C Vaucher, Philippe Schwaller, and Teodoro Laino. 2023 · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023b · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Survey of Protein Sequence Embedding Models
Chau Tran, Siddharth Khadkikar, and Aleksey Porollo. 2023 · 2023
Later among the works it cites.
Molecular Descriptors Property Prediction Using Transformer-Based Approach
Tuan Tran and Chinwe Ekenna. 2023 · 2023
Later among the works it cites.
Deciphering the protein landscape with ProtFlash: a lightweight language model
Lei Wang, Hui Zhang, Wei Xu, Zhidong Xue, and Yan Wang. 2023i · 2023
Later among the works it cites.
miProBERT: identification of microRNA promoters based on the pre-trained model BERT
Xin Wang, Xin Gao, Guohua Wang, and Dan Li. 2023a · 2023
Later among the works it cites.
UNI-RNA: universal pre-trained models revolutionize RNA research
Xi Wang, Ruichu Gu, Zhiyuan Chen, Yongge Li, Xiaohong Ji, Guolin Ke, and Han Wen. 2023b · 2023
Later among the works it cites.
cMolGPT: A Conditional Generative Pre-Trained Transformer for Target-Specific De Novo Molecular Generation
Ye Wang, Honggang Zhao, Simone Sciabola, and Wenlu Wang. 2023j · 2023
Later among the works it cites.
BioBridge: Bridging Biomedical Foundation Models via Knowledge Graph
Zifeng Wang, Zichen Wang, Balasubramaniam Srinivasan, Vassilis N Ioannidis, Huzefa Rangwala, and Rishita Anubhai. 2023e · 2023
Later among the works it cites.
InstructProtein: Aligning Human and Protein Language via Knowledge Instruction
Zeyuan Wang, Qiang Zhang, Keyan Ding, Ming Qin, Xiang Zhuang, Xiaotong Li, and Huajun Chen. 2023g · 2023
Later among the works it cites.
A Systematic Survey of Chemical Pre-trained Models. IJCAI
Jun Xia, Yanqiao Zhu, Yuanqi Du, Y Liu, and SZ Li. 2023 · 2023
Later among the works it cites.
DARWIN Series: Domain Specific Large Language Models for Natural Science
Tong Xie, Yuwei Wan, Wei Huang, Zhenyu Yin, Yixuan Liu, Shaozhou Wang, Qingyuan Linghu, Chunyu Kit, Clara Grazian, Wenjie Zhang, Imran Razzak, and Bram Hoex. 2023 · 2023
Later among the works it cites.
DoctorGLM: Fine-tuning your Chinese Doctor is not a Herculean Task
Honglin Xiong, Sheng Wang, Yitao Zhu, Zihao Zhao, Yuxiao Liu, Linlin Huang, Qian Wang, and Dinggang Shen. 2023 · 2023
Later among the works it cites.
Baize: An open-source chat model with parameter-efficient tuning on self-chat data
Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2023a · 2023
Later among the works it cites.
Multilingual translation for zero-shot biomedical classification using BioTranslator
Hanwen Xu, Addie Woicik, Hoifung Poon, Russ B Altman, and Sheng Wang. 2023b · 2023
Later among the works it cites.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Bing Yin, and Xia Hu. 2023a · 2023
Later among the works it cites.
Qilin-Med: Multi-stage Knowledge Injection Advanced Medical Large Language Model
Qichen Ye, Junling Liu, Dading Chong, Peilin Zhou, Yining Hua, and Andrew Liu. 2023 · 2023
Later among the works it cites.
Selformer: Molecular representation learning via selfies language models
Atakan Yüksel, Erva Ulusoy, Atabey Ünlü, and Tunca Doğan. 2023 · 2023
Later among the works it cites.
The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods
Barbara Zdrazil, Eloy Felix, Fiona Hunter, Emma J Manners, James Blackshaw, Sybilla Corbett, Marleen de Veij, Harris Ioannidis, David Mendez Lopez, Juan F Mosquera, et al · 2023
Later among the works it cites.
DNAGPT: A Generalized Pretrained Tool for Multiple DNA Sequence Analysis Tasks
Daoan Zhang, Weitong Zhang, Bing He, Jianguo Zhang, Chenchen Qin, and Jianhua Yao. 2023f · 2023
Later among the works it cites.
HuatuoGPT, towards Taming Language Model to Be a Doctor
Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Jianquan Li, Guiming Chen, Xiangbo Wu, Zhiyi Zhang, Qingying Xiao, Xiang Wan, Benyou Wang, and Haizhou Li. 2023 · 2023
Later among the works it cites.
Enhancing the Protein Tertiary Structure Prediction by Multiple Sequence Alignment Generation
Le Zhang, Jiayang Chen, Tao Shen, Yu Li, and Siqi Sun. 2023 · 2023
Later among the works it cites.
Prediction of multiple types of RNA modifications via biological language model
Ying Zhang, Fang Ge, Fuyi Li, Xibei Yang, Jiangning Song, and Dong-Jun Yu. 2023a · 2023
Later among the works it cites.
Multiple sequence-alignment-based RNA language model and its application to structural inference
Yikun Zhang, Mei Lang, Jiuhong Jiang, Zhiqiang Gao, Fan Xu, Thomas Litfin, Ke Chen, Jaswinder Singh, Xiansong Huang, Guoli Song, et al · 2023
Later among the works it cites.
Enhancing protein language models with structure-based encoder and pre-training
Zuobai Zhang, Minghao Xu, Vijil Chenthamarakshan, Aurélie Lozano, Payel Das, and Jian Tang. 2023e · 2023
Later among the works it cites.
GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot Learning
Haiteng Zhao, Shengchao Liu, Chang Ma, Hannan Xu, Jie Fu, Zhi-Hong Deng, Lingpeng Kong, and Qi Liu. 2023a · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Structure-informed Language Models Are Protein Designers
Zaixiang Zheng, Yifan Deng, Dongyu Xue, Yi Zhou, Fei Ye, and Quanquan Gu. 2023a · 2023
Later among the works it cites.
ChatGPT Chemistry Assistant for Text Mining and the Prediction of MOF Synthesis
Zhiling Zheng, Oufan Zhang, Christian Borgs, Jennifer T. Chayes, and Omar M. Yaghi. 2023b · 2023
Later among the works it cites.
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023 · 2023
Later among the works it cites.
Uni-Mol: a universal 3D molecular representation learning framework
Gengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng, Hongteng Xu, Zhewei Wei, Linfeng Zhang, and Guolin Ke. 2023a · 2023
Later among the works it cites.
Protein Representation Learning via Knowledge Enhanced Primary Structure Modeling
Hong-Yu Zhou, Yunxiang Fu, Zhicheng Zhang, Cheng Bian, and Yizhou Yu. 2023 · 2023
Later among the works it cites.
Dnabert-2: Efficient foundation model and benchmark for multi-species genome
Zhihan Zhou, Yanrong Ji, Weijian Li, Pratik Dutta, Ramana Davuluri, and Han Liu. 2023b · 2023
Later among the works it cites.
Learning Over Molecular Conformer Ensembles: Datasets and Benchmarks
Yanqiao Zhu, Jeehyun Hwang, Keir Adams, Zhen Liu, Bozhao Nan, Brock Stenfors, Yuanqi Du, Jatin Chauhan, Olaf Wiest, Olexandr Isayev, et al · 2023
Later among the works it cites.
INDUS: Effective and Efficient Language Models for Scientific Applications
Bishwaranjan Bhattacharjee, Aashka Trivedi, Masayasu Muraoka, Muthukumaran Ramasubramanian, Takuma Udagawa, Iksha Gurung, Rong Zhang, Bharath Dandala, Rahul Ramachandran, Manil Maskey, et al · 2024
Closest in time.
Sciassess: Benchmarking llm proficiency in scientific literature analysis
Hengxing Cai, Xiaochen Cai, Junhan Chang, Sihang Li, Lin Yao, Changxin Wang, Zhifeng Gao, Yongge Li, Mujie Lin, Shuwen Yang, et al · 2024
Closest in time.
Uni-SMART: Universal Science Multimodal Analysis and Research Transformer
Hengxing Cai, Xiaochen Cai, Shuwen Yang, Jiankun Wang, Lin Yao, Zhifeng Gao, Junhan Chang, Sihang Li, Mingjun Xu, Changxin Wang, et al · 2024
Closest in time.
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, et al · 2024
Closest in time.
Bidirectional generation of structure and properties through a single molecular foundation model
Jinho Chang and Jong Chul Ye. 2024 · 2024
Closest in time.
PharmGPT: Domain-Specific Large Language Models for Bio-Pharmaceutical and Chemistry
Linqing Chen, Weilei Wang, Zilong Bai, Peng Xu, Yan Fang, Jie Fang, Wentao Wu, Lizhi Zhou, Ruiji Zhang, Yubin Xia, et al · 2024
Closest in time.
FGBERT: Function-Driven Pre-trained Gene Language Model for Metagenomics
ChenRui Duan, Zelin Zang, Yongjie Xu, Hang He, Zihan Liu, Zijia Song, Ju-Sheng Zheng, and Stan Z. Li. 2024 · 2024
Closest in time.
Evaluating generalizability of artificial intelligence models for molecular datasets
Yasha Ektefaie, Andrew Shen, Daria Bykova, Maximillian Marin, Marinka Zitnik, and Maha R Farhat. 2024 · 2024
Closest in time.
SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models
Kehua Feng, Keyan Ding, Weijie Wang, Xiang Zhuang, Zeyuan Wang, Ming Qin, Yu Zhao, Jianhua Yao, Qiang Zhang, and Huajun Chen. 2024 · 2024
Closest in time.
Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain
Iker García-Ferrero, Rodrigo Agerri, Aitziber Atutxa Salazar, Elena Cabrio, Iker de la Iglesia, Alberto Lavelli, Bernardo Magnini, Benjamin Molinet, Johana Ramirez-Romero, German Rigau, Jose Maria Villa-Gonzalez, Serena Villata, and Andrea Zaninello. 2024 · 2024
Closest in time.
Simulating 500 million years of evolution with a language model
Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J. Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q. Tran, Jonathan Deaton, Marius Wiggert, Rohil Badkundri, Irhum Shafkat, Jun Gong, Alexander Derry, Raul S. Molina, Neil Thomas, Yousuf Khan, Chetan Mishra, Carolyn Kim, Liam J. Bartie, Matthew Nemeth, Patrick D. Hsu, Tom Sercu, Salvatore Candido, and Alexander Rives. 2024 · 2024
Closest in time.
Assessing GPT-4 for cell type annotation in single-cell RNA-seq analysis
Wenpin Hou and Zhicheng Ji. 2024 · 2024
Closest in time.
Genomic language model predicts protein co-regulation and function
Yunha Hwang, Andre L Cornman, Elizabeth H Kellogg, Sergey Ovchinnikov, and Peter R Girguis. 2024 · 2024
Closest in time.
Small Language Models Learn Enhanced Reasoning Skills from Medical Textbooks
Hyunjae Kim, Hyeon Hwang, Jiwoo Lee, Sihyeon Park, Dain Kim, Taewhoo Lee, Chanwoong Yoon, Jiwoong Sohn, Donghee Choi, and Jaewoo Kang. 2024 · 2024
Closest in time.
Biomistral: A collection of open-source pretrained large language models for medical domains
Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024 · 2024
Closest in time.
Exploring Genomic Large Language Models: Bridging the Gap between Natural Language and Gene Sequences
Huaqing Liu, Shuxian Zhou, Peiyi Chen, Jiahui Liu, Ku-Geng Huo, and Lanqing Han. 2024e · 2024
Closest in time.
DrugLLM: Open Large Language Model for Few-shot Molecule Generation
Xianggen Liu, Yan Guo, Haoran Li, Jin Liu, Shudong Huang, Bowen Ke, and Jiancheng Lv. 2024b · 2024
Closest in time.
MolecularGPT: Open Large Language Model (LLM) for Few-Shot Molecular Property Prediction
Yuyan Liu, Sirui Ding, Sheng Zhou, Wenqi Fan, and Qiaoyu Tan. 2024a · 2024
Closest in time.
GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models
Zicheng Liu, Jiahui Li, Siyuan Li, Zelin Zang, Cheng Tan, Yufei Huang, Yajing Bai, and Stan Z Li. 2024c · 2024
Closest in time.
ProtT3: Protein-to-Text Generation for Text-based Protein Understanding
Zhiyuan Liu, An Zhang, Hao Fei, Enzhi Zhang, Xiang Wang, Kenji Kawaguchi, and Tat-Seng Chua. 2024d · 2024
Closest in time.
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension
Xingyu Lu, He Cao, Zijing Liu, Shengyuan Bai, Leqing Chen, Yuan Yao, Hai-Tao Zheng, and Yu Li. 2024 · 2024
Closest in time.
ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing
Liuzhenghao Lv, Zongying Lin, Hao Li, Yuyang Liu, Jiaxi Cui, Calvin Yu-Chian Chen, Li Yuan, and Yonghong Tian. 2024 · 2024
Closest in time.
Relative molecule self-attention transformer
Łukasz Maziarka, Dawid Majchrowski, Tomasz Danel, Piotr Gaiński, Jacek Tabor, Igor Podolak, Paweł Morkisz, and Stanisław Jastrzębski. 2024 · 2024
Closest in time.
Sequence modeling and design from molecular to genome scale with Evo
Eric Nguyen, Michael Poli, Matthew G Durrant, Armin W Thomas, Brian Kang, Jeremy Sullivan, Madelena Y Ng, Ashley Lewis, Aman Patel, Aaron Lou, et al · 2024
Closest in time.
Codon language embeddings provide strong signals for use in protein engineering
Carlos Outeiral and Charlotte M Deane. 2024 · 2024
Closest in time.
Biot5+: Towards generalized biological understanding with iupac integration and multi-task tuning
Qizhi Pei, Lijun Wu, Kaiyuan Gao, Xiaozhuan Liang, Yin Fang, Jinhua Zhu, Shufang Xie, Tao Qin, and Rui Yan. 2024a · 2024
Closest in time.
Leveraging Biomolecule and Natural Language through Multi-Modal Learning: A Survey
Qizhi Pei, Lijun Wu, Kaiyuan Gao, Jinhua Zhu, Yue Wang, Zun Wang, Tao Qin, and Rui Yan. 2024c · 2024
Closest in time.
3D-MolT5: Towards Unified 3D Molecule-Text Modeling with 3D Molecular Tokenization
Qizhi Pei, Lijun Wu, Kaiyuan Gao, Jinhua Zhu, and Rui Yan. 2024b · 2024
Closest in time.
Rinalmo: General-purpose rna language models can generalize well on structure prediction tasks
Rafael Josip Penić, Tin Vlašić, Roland G Huber, Yue Wan, and Mile Šikić. 2024 · 2024
Closest in time.
BiMediX: Bilingual Medical Mixture of Experts LLM
Sara Pieri, Sahal Shaji Mullappilly, Fahad Shahbaz Khan, Rao Muhammad Anwer, Salman Khan, Timothy Baldwin, and Hisham Cholakkal. 2024 · 2024
Closest in time.
A Review of Large Language Models and Autonomous Agents in Chemistry
Mayk Caldas Ramos, Christopher J. Collison, and Andrew D. White. 2024 · 2024
Closest in time.
BEACON: Benchmark for Comprehensive RNA Tasks and Language Models
Yuchen Ren, Zhiyuan Chen, Lifeng Qiao, Hongtai Jing, Yuchen Cai, Sheng Xu, Peng Ye, Xinzhu Ma, Siqi Sun, Hongliang Yan, et al · 2024
Closest in time.
Joint Embedding of Transcriptomes and Text Enables Interactive Single-Cell RNA-seq Data Exploration via Natural Language. In ICLR 2024 Workshop on Machine Learning for Genomics Explorations
Moritz Schaefer, Peter Peneder, Daniel Malzl, Anna Hakobyan, Varun S Sharma, Thomas Krausgruber, Jörg Menche, Eleni Tomazou, and Christoph Bock. [n. d.] · 2024
Closest in time.
Caduceus: Bi-directional equivariant long-range dna sequence modeling
Yair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao, Albert Gu, and Volodymyr Kuleshov. 2024 · 2024
Closest in time.
A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding
Yiqing Shen, Zan Chen, Michail Mamalakis, Luhan He, Haiyang Xia, Tianbin Li, Yanzhou Su, Junjun He, and Yu Guang Wang. 2024 · 2024
Closest in time.
Multi-purpose RNA language modelling with motif-aware pretraining and type-guided fine-tuning
Ning Wang, Jiang Bian, Yuchen Li, Xuhong Li, Shahid Mumtaz, Linghe Kong, and Haoyi Xiong. 2024 · 2024
Closest in time.
Me LLaMA: Foundation Large Language Models for Medical Applications
Qianqian Xie, Qingyu Chen, Aokun Chen, Cheng Peng, Yan Hu, Fongci Lin, Xueqing Peng, Jimin Huang, Jeffrey Zhang, Vipina Keloth, Xinyu Zhou, Huan He, Lucila Ohno-Machado, Yonghui Wu, Hua Xu, and Jiang Bian. 2024 · 2024
Closest in time.
Botao Yu, Frazier N Baker, Ziqi Chen, Xia Ning, and Huan Sun. 2024 · 2024
Closest in time.
Functional Protein Design with Local Domain Alignment
Chaohao Yuan, Songyou Li, Geyan Ye, Yikun Zhang, Long-Kai Huang, Wenbing Huang, Wei Liu, Jianhua Yao, and Yu Rong. 2024 · 2024
Closest in time.
Sciglm: Training scientific language models with self-reflective instruction annotation and tuning
Dan Zhang, Ziniu Hu, Sining Zhoubian, Zhengxiao Du, Kaiyu Yang, Zihan Wang, Yisong Yue, Yuxiao Dong, and Jie Tang. 2024a · 2024
Closest in time.
Chemllm: A chemical large language model
Di Zhang, Wei Liu, Qian Tan, Jingdan Chen, Hang Yan, Yuliang Yan, Jiatong Li, Weiran Huang, Xiangyu Yue, Dongzhan Zhou, et al · 2024
Closest in time.
Atomas: Hierarchical Alignment on Molecule-Text for Unified Molecule Understanding and Generation
Yikun Zhang, Geyan Ye, Chaohao Yuan, Bo Han, Long-Kai Huang, Jianhua Yao, Wei Liu, and Yu Rong. 2024c · 2024
Closest in time.
LangCell: Language-Cell Pre-training for Cell Identity Understanding
Suyuan Zhao, Jiahuan Zhang, Yizhen Luo, Yushuai Wu, and Zaiqing Nie. 2024b · 2024
Closest in time.
Chemdfm: Dialogue foundation model for chemistry
Zihan Zhao, Da Ma, Lu Chen, Liangtai Sun, Zihao Li, Hongshen Xu, Zichen Zhu, Su Zhu, Shuai Fan, Guodong Shen, et al · 2024
Closest in time.
Multi-Scale Protein Language Model for Unified Molecular Modeling
Kangjie Zheng, Siyu Long, Tianyu Lu, Junwei Yang, Xinyu Dai, Ming Zhang, Zaiqing Nie, Wei-Ying Ma, and Hao Zhou. 2024 · 2024
Closest in time.
Protllm: An interleaved protein-language llm with protein-as-word pre-training
Le Zhuo, Zewen Chi, Minghao Xu, Heyan Huang, Heqi Zheng, Conghui He, Xian-Ling Mao, and Wentao Zhang. 2024 · 2024
Closest in time.
Automated chemical reaction extraction from scientific literature
Jiang Guo, A Santiago Ibanez-Lopez, Hanyu Gao, Victor Quach, Connor W Coley, Klavs F Jensen, and Regina Barzilay. 2021 · 2045
Closest in time.
MolGPT: molecular generation using a transformer-decoder model
Viraj Bagal, Rishal Aggarwal, PK Vinod, and U Deva Priyakumar. 2021 · 2076
Closest in time.