Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have revolutionized the field of natural language processing, but they fall short in comprehending biological sequences such as proteins.
Nonmetric multidimensional scaling: a numerical method
Joseph B Kruskal · 1964
Earlier work this paper cites.
The synaptic vesicle cycle: a cascade of protein–protein interactions
Thomas C Südhof · 1995
Earlier work this paper cites.
Scaffolding proteins organize multimolecular protein complexes for sensory signal transduction
Armin Huber · 2001
Earlier work this paper cites.
The integrated analysis of metabolic and protein interaction networks reveals novel molecular organizing principles
Pawel Durek and Dirk Walther · 2008
Earlier work this paper cites.
AutoDock Vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading
Oleg Trott and Arthur J Olson · 2010
Earlier work this paper cites.
Lessons learned in empirical scoring with smina from the csar 2011 benchmarking exercise
David Ryan Koes, Matthew P Baumgartner, and Carlos J Camacho · 2013
Earlier work this paper cites.
UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches
Baris E Suzek, Yuqi Wang, Hongzhan Huang, Peter B McGarvey, Cathy H Wu, and UniProt Consortium · 2015
Earlier work this paper cites.
Protein regulation in signal transduction
Michael J Lee and Michael B Yaffe · 2016
Earlier work this paper cites.
Deeploc: prediction of protein subcellular localization using deep learning
José Juan Almagro Armenteros, Casper Kaae Sønderby, Søren Kaae Sønderby, Henrik Nielsen, and Ole Winther · 2017
Earlier work this paper cites.
The most popular genes in the human genome
Elie Dolgin · 2017
Earlier work this paper cites.
MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets
Martin Steinegger and Johannes Söding · 2017
Earlier work this paper cites.
Deepsf: deep convolutional neural network for mapping protein sequences to folds
Jie Hou, Badri Adhikari, and Jianlin Cheng · 2018
Earlier work this paper cites.
Deep Contextualized Word Representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Unified rational protein engineering with sequence-based deep representation learning
Ethan C Alley, Grigory Khimulya, Surojit Biswas, Mohammed AlQuraishi, and George M Church · 2019
Earlier work this paper cites.
UniProt: a worldwide hub of protein knowledge
UniProt Consortium · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Learning the Difference That Makes A Difference with Counterfactually-Augmented Data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton · 2019
Earlier work this paper cites.
Do Not Have Enough Data? Deep Learning to the Rescue!
Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, Naama Tepper, and Naama Zwerdling · 2020
Earlier work this paper cites.
Good-Enough Compositional Data Augmentation
Jacob Andreas · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, and others Askell · 2020
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
K-BERT: Enabling Language Representation with Knowledge Graph
Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Qi Ju, Haotang Deng, and Ping Wang · 2020
Earlier work this paper cites.
Transformer Protein Language Models Are Unsupervised Structure Learners
Roshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov, and Alexander Rives · 2020
Earlier work this paper cites.
CoLAKE: Contextualized Language and Knowledge Embedding
Tianxiang Sun, Yunfan Shao, Xipeng Qiu, Qipeng Guo, Yaru Hu, Xuan-Jing Huang, and Zheng Zhang · 2020
Cited alongside, same era.
Learning from Task Descriptions
Orion Weller, Nicholas Lourie, Matt Gardner, and Matthew E Peters · 2020
Cited alongside, same era.
Evaluating Large Language Models Trained on Code, 2021
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Cited alongside, same era.
ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning
Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, et al · 2021
Cited alongside, same era.
Structure-based protein function prediction using graph convolutional networks
Vladimir Gligorijević, P Douglas Renfrew, Tomasz Kosciolek, Julia Koehler Leman, Daniel Berenberg, Tommi Vatanen, Chris Chandler, Bryn C Taylor, Ian M Fisk, Hera Vlamakis, et al · 2021
Cited alongside, same era.
An open invitation to the Understudied Proteins Initiative
Georg Kustatscher, Tom Collins, Anne-Claude Gingras, Tiannan Guo, Henning Hermjakob, Trey Ideker, Kathryn S Lilley, Emma Lundberg, Edward M Marcotte, Markus Ralser, et al · 2022
Later among the works it cites.
ColabFold: Making Protein folding accessible to all
Milot Mirdita, Konstantin Schütze, Yoshitaka Moriwaki, Lim Heo, Sergey Ovchinnikov, and Martin Steinegger · 2022
Later among the works it cites.
Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi · 2022
Later among the works it cites.
Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval
Pascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena Hurtado, Aidan N Gomez, Debora Marks, and Yarin Gal · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits · 2021
Cited alongside, same era.
Highly accurate protein structure prediction with AlphaFold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Cited alongside, same era.
Global mapping of protein–metabolite interactions in saccharomyces cerevisiae reveals that ser-leu dipeptide regulates phosphoglycerate kinase activity
Marcin Luzarowski, Rubén Vicente, Andrei Kiselev, Mateusz Wagner, Dennis Schlossarek, Alexander Erban, Leonardo Perez de Souza, Dorothee Childs, Izabela Wojciechowska, Urszula Luzarowska, et al · 2021
Cited alongside, same era.
Language models enable zero-shot prediction of the effects of mutations on protein function
Joshua Meier, Roshan Rao, Robert Verkuil, Jason Liu, Tom Sercu, and Alex Rives · 2021
Cited alongside, same era.
Scaling Language Models: Methods, Analysis & Insights from Training Gopher, 2021
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Cited alongside, same era.
MSA Transformer
Roshan M Rao, Jason Liu, Robert Verkuil, Joshua Meier, John Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives · 2021
Cited alongside, same era.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al · 2021
Cited alongside, same era.
MedMCQA: A Large-scale Multi-subject Multi-Choice Dataset for Medical domain Question Answering
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu · 2022
Later among the works it cites.
A Generalist Agent
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gómez Colmenarejo, Alexander Novikov, Gabriel Barth-maron, Mai Giménez, Yury Sulsky, Jackie Kay, et al · 2022
Later among the works it cites.
Galactica: A Large Language Model for Science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic · 2022
Later among the works it cites.
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, et al · 2022
Later among the works it cites.
ZeroGen: Efficient Zero-shot Learning via Dataset Generation
Jiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu, Jiangtao Feng, Zhiyong Wu, Tao Yu, and Lingpeng Kong · 2022
Later among the works it cites.
Prot2Text: Multimodal Protein’s Function Generation with GNNs and Transformers
Hadi Abdine, Michail Chatzianastasis, Costas Bouyioukos, and Michalis Vazirgiannis · 2023
Closest in time.
The Gene Ontology knowledgebase in 2023
Suzi A Aleksander, James Balhoff, Seth Carbon, J Michael Cherry, Harold J Drabkin, Dustin Ebert, Marc Feuermann, Pascale Gaudet, Nomi L Harris, et al · 2023
Closest in time.
Diffdock: Diffusion Steps, Twists, and Turns for Molecular Docking
Gabriele Corso, Hannes Stärk, Bowen Jing, Regina Barzilay, and Tommi Jaakkola · 2023
Closest in time.
MultiModal-GPT: A Vision and Language Model for Dialogue with Humans, 2023
Tao Gong, Chengqi Lyu, Shilong Zhang, Yudong Wang, Miao Zheng, Qian Zhao, Kuikun Liu, Wenwei Zhang, Ping Luo, and Kai Chen · 2023
Closest in time.
Unnatural instructions: Tuning language models with (almost) no human labor
Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick · 2023
Closest in time.
Language Is Not All You Need: Aligning Perception with Language Models, 2023
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, et al · 2023
Closest in time.
Evolutionary-scale prediction of atomic-level protein structure with a language model
Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al · 2023
Closest in time.
BioMedGPT: Open Multimodal Generative Pre-trained Transformer for Biomedicine
Yizhen Luo, Jiahuan Zhang, Siqi Fan, Kai Yang, Yushuai Wu, Mu Qiao, and Zaiqing Nie · 2023
Closest in time.
InterPro in 2022
Typhaine Paysan-Lafosse, Matthias Blum, Sara Chuguransky, Tiago Grego, Beatriz Lázaro Pinto, Gustavo A Salazar, Maxwell L Bileschi, Peer Bork, Alan Bridge, Lucy Colwell, et al · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Closest in time.
Protst: Multi-Modality Learning of Protein Sequences and Biomedical Texts
Minghao Xu, Xinyu Yuan, Santiago Miret, and Jian Tang · 2023
Closest in time.