Fetching the paper…
Reading the bibliography…
In many scientific fields, large language models (LLMs) have revolutionized the way text and other modalities of data (e.g., molecules and proteins) are handled, achieving superior performance in various applications and augmenting the scientific discovery process.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Clinicalbert: Modeling clinical notes and predicting hospital readmission
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. 2019 · 1904
Earlier work this paper cites.
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. 2019 · 1907
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Clinical concept extraction using transformers
Xi Yang, Jiang Bian, William R Hogan, and Yonghui Wu. 2020 · 1942
Earlier work this paper cites.
Overview of the coupled model intercomparison project phase 6 (cmip6) experimental design and organization
Veronika Eyring, Sandrine Bony, Gerald A Meehl, Catherine A Senior, Bjorn Stevens, Ronald J Stouffer, and Karl E Taylor. 2016 · 1958
Earlier work this paper cites.
Rapid and sensitive protein similarity searches
David J Lipman and William R Pearson. 1985 · 1985
Earlier work this paper cites.
An introduction to wu’s method for mechanical theorem proving in geometry
Shang-Ching Chou. 1988 · 1988
Earlier work this paper cites.
Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules
David Weininger. 1988 · 1988
Earlier work this paper cites.
The swiss-prot protein sequence database and its supplement trembl in 2000
Amos Bairoch and Rolf Apweiler. 2000 · 2000
Earlier work this paper cites.
Molecule attention transformer
Łukasz Maziarka, Tomasz Danel, Sławomir Mucha, Krzysztof Rataj, Jacek Tabor, and Stanisław Jastrzębski. 2020 · 2002
Earlier work this paper cites.
Pubmed central (pmc): An archive for literature from life sciences journals
Jeff Beck and Ed Sequeira. 2003 · 2003
Earlier work this paper cites.
The unified medical language system (umls): integrating biomedical terminology
Olivier Bodenreider. 2004 · 2004
Earlier work this paper cites.
Pre-training technique to localize medical bert and enhance biomedical bert
Shoya Wada, Toshihiro Takeda, Shiro Manabe, Shozo Konishi, Jun Kamohara, and Yasushi Matsumura. 2020 · 2005
Earlier work this paper cites.
Openstreetmap: User-generated street maps
Mordechai Haklay and Patrick Weber. 2008 · 2008
Earlier work this paper cites.
Arnetminer: extraction and mining of academic social networks
Jie Tang, Jing Zhang, Limin Yao, Juanzi Li, Li Zhang, and Zhong Su. 2008 · 2008
Earlier work this paper cites.
Conceptualized representation learning for chinese biomedical text mining
Ningyu Zhang, Qianghuai Jia, Kangping Yin, Liang Dong, Feng Gao, and Nengwei Hua. 2020 · 2008
Earlier work this paper cites.
Ape210k: A large-scale and template-rich dataset of math word problems
Wei Zhao, Mingyue Shang, Yang Liu, Liang Wang, and Jingming Liu. 2020 · 2009
Earlier work this paper cites.
Chemberta: large-scale self-supervised pretraining for molecular property prediction
Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. 2020 · 2010
Earlier work this paper cites.
Molecular representation learning with language models and domain-relevant auxiliary tasks
Benedek Fabian, Thomas Edlich, Héléna Gaspar, Marwin Segler, Joshua Meyers, Marco Fiscato, and Mohamed Ahmed. 2020 · 2011
Earlier work this paper cites.
Pubmed and beyond: a survey of web tools for searching biomedical literature
Zhiyong Lu. 2011 · 2011
Earlier work this paper cites.
Gencode: the reference human genome annotation for the encode project
Jennifer Harrow, Adam Frankish, Jose M Gonzalez, Electra Tapanari, Mark Diekhans, Felix Kokocinski, Bronwen L Aken, Daniel Barrell, Amonida Zadissa, Stephen Searle, et al. 2012 · 2012
Earlier work this paper cites.
Commentary: The materials project: A materials genome approach to accelerating materials innovation
Anubhav Jain, Shyue Ping Ong, Geoffroy Hautier, Wei Chen, William Davidson Richards, Stephen Dacek, Shreyas Cholia, Dan Gunter, David Skinner, Gerbrand Ceder, et al. 2013 · 2013
Earlier work this paper cites.
Open-source platform to benchmark fingerprints for ligand-based virtual screening
Sereina Riniker and Gregory A Landrum. 2013 · 2013
Earlier work this paper cites.
Ncbi disease corpus: a resource for disease name recognition and concept normalization
Rezarta Islamaj Doğan, Robert Leaman, and Zhiyong Lu. 2014 · 2014
Earlier work this paper cites.
Quantum chemistry structures and properties of 134 kilo molecules
Raghunathan Ramakrishnan, Pavlo O Dral, Matthias Rupp, and O Anatole Von Lilienfeld. 2014 · 2014
Earlier work this paper cites.
A global reference for human genetic variation
The 1000 Genomes Project Consortium. 2015 · 2015
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
Panupong Pasupat and Percy Liang. 2015 · 2015
Earlier work this paper cites.
Solving geometry problems: Combining text and diagram interpretation
Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, Oren Etzioni, and Clint Malcolm. 2015 · 2015
Earlier work this paper cites.
An overview of microsoft academic service (mas) and applications
Arnab Sinha, Zhihong Shen, Yang Song, Hao Ma, Darrin Eide, Bo-June Hsu, and Kuansan Wang. 2015 · 2015
Earlier work this paper cites.
Zinc 15–ligand discovery for everyone
Teague Sterling and John J Irwin. 2015 · 2015
Earlier work this paper cites.
Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches
Baris E Suzek, Yuqi Wang, Hongzhan Huang, Peter B McGarvey, Cathy H Wu, and UniProt Consortium. 2015 · 2015
Earlier work this paper cites.
A full-text learning to rank dataset for medical information retrieval
Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016 · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016 · 2016
Earlier work this paper cites.
A large public corpus of web tables containing time and context metadata
Oliver Lehmberg, Dominique Ritze, Robert Meusel, and Christian Bizer. 2016 · 2016
Earlier work this paper cites.
What’s what: The (nearly) definitive guide to reaction role assignment
Nadine Schneider, Nikolaus Stiefl, and Gregory A Landrum. 2016 · 2016
Earlier work this paper cites.
The chembl database in 2017
Anna Gaulton, Anne Hersey, Michał Nowotka, A Patricia Bento, Jon Chambers, David Mendez, Prudence Mutowo, Francis Atkinson, Louisa J Bellis, Elena Cibrián-Uhalte, et al. 2017 · 2017
Earlier work this paper cites.
Predicting organic reaction outcomes with weisfeiler-lehman network
Wengong Jin, Connor W Coley, Regina Barzilay, and Tommi Jaakkola. 2017 · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling. 2017 · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. 2017 · 2017
Earlier work this paper cites.
Deep neural solver for math word problems
Yan Wang, Xiaojiang Liu, and Shuming Shi. 2017 · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F Liu, and Matt Gardner. 2017 · 2017
Earlier work this paper cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Victor Zhong, Caiming Xiong, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Construction of the literature graph in semantic scholar
Waleed Ammar, Dirk Groeneveld, Chandra Bhagavatula, Iz Beltagy, Miles Crawford, Doug Downey, Jason Dunkelberger, Ahmed Elgohary, Sergey Feldman, Vu Ha, et al. 2018 · 2018
Earlier work this paper cites.
Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction
Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018 · 2018
Earlier work this paper cites.
Radiology objects in context (roco): a multimodal image dataset
Obioma Pelka, Sven Koitka, Johannes Rückert, Felix Nensa, and Christoph M Friedrich. 2018 · 2018
Earlier work this paper cites.
Lessons from natural language inference in the clinical domain
Alexey Romanov and Chaitanya Shivade. 2018 · 2018
Earlier work this paper cites.
Moleculenet: a benchmark for molecular machine learning
Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. 2018 · 2018
Earlier work this paper cites.
Publicly available clinical bert embeddings
Emily Alsentzer, John Murphy, William Boag, Wei-Hung Weng, Di Jindi, Tristan Naumann, and Matthew McDermott. 2019 · 2019
Earlier work this paper cites.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019 · 2019
Earlier work this paper cites.
Scibert: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Earlier work this paper cites.
Structural scaffolds for citation intent classification in scientific publications
Arman Cohan, Waleed Ammar, Madeleine van Zuylen, and Field Cady. 2019 · 2019
Earlier work this paper cites.
Rnacentral: a hub of information for non-coding rna sequences
The RNAcentral Consortium. 2019 · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Panglaodb: a web server for exploration of mouse and human single-cell rna sequencing data
Oscar Franzén, Li-Ming Gan, and Johan LM Björkegren. 2019 · 2019
Earlier work this paper cites.
Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison
Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Silviana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, et al. 2019 · 2019
Earlier work this paper cites.
Probing biomedical embeddings from language models
Qiao Jin, Bhuwan Dhingra, William Cohen, and Xinghua Lu. 2019 · 2019
Earlier work this paper cites.
Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports
Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng. 2019 · 2019
Earlier work this paper cites.
Pubchem 2019 update: improved access to chemical data
Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al. 2019 · 2019
Earlier work this paper cites.
Fine-tuning bidirectional encoder representations from transformers (bert)–based models on large-scale electronic health record notes: an empirical study
Fei Li, Yonghao Jin, Weisong Liu, Bhanu Pratap Singh Rawat, Pengshan Cai, Hong Yu, et al. 2019 · 2019
Earlier work this paper cites.
Transfer learning in biomedical natural language processing: An evaluation of bert and elmo on ten benchmarking datasets
Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019 · 2019
Earlier work this paper cites.
Pre-training of graph augmented transformers for medication recommendation
Junyuan Shang, Tengfei Ma, Cao Xiao, and Jimeng Sun. 2019 · 2019
Earlier work this paper cites.
Smiles-bert: large scale unsupervised pre-training for molecular property prediction
Sheng Wang, Yuzhi Guo, Yuhong Wang, Hongmao Sun, and Junzhou Huang. 2019 · 2019
Earlier work this paper cites.
Cometa: A corpus for medical entity linking in the social media
Marco Basaldella, Fangyu Liu, Ehsan Shareghi, and Nigel Collier. 2020 · 2020
Earlier work this paper cites.
Highly accurate classification of chest radiographic reports using a deep learning natural language model pre-trained on 3.8 million text reports
Keno K Bressem, Lisa C Adams, Robert A Gaudin, Daniel Tröltzsch, Bernd Hamm, Marcus R Makowski, Chan-Yong Schüle, Janis L Vahldiek, and Stefan M Niehues. 2020 · 2020
Earlier work this paper cites.
Padchest: A large chest x-ray image dataset with multi-label annotated reports
Aurelia Bustos, Antonio Pertusa, Jose-Maria Salinas, and Maria De La Iglesia-Vaya. 2020 · 2020
Earlier work this paper cites.
Tldr: Extreme summarization of scientific documents
Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel S Weld. 2020 · 2020
Earlier work this paper cites.
Biomedbert: A pre-trained biomedical language model for qa and ir
Souradip Chakraborty, Ekaba Bisong, Shweta Bhatt, Thomas Wagner, Riley Elliott, and Francesco Mosconi. 2020 · 2020
Earlier work this paper cites.
Specter: Document-level representation learning using citation-informed transformers
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S Weld. 2020 · 2020
Earlier work this paper cites.
Injecting numerical reasoning skills into language models
Mor Geva, Ankit Gupta, and Jonathan Berant. 2020 · 2020
Earlier work this paper cites.
Tapas: Weakly supervised table parsing via pre-training
Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Mueller, Francesco Piccinno, and Julian Eisenschlos. 2020 · 2020
Earlier work this paper cites.
Clinical xlnet: Modeling sequential clinical notes and predicting prolonged mechanical ventilation
Kexin Huang, Abhishek Singh, Sitong Chen, Edward Moseley, Chih-Ying Deng, Naomi George, and Charolotta Lindvall. 2020 · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Earlier work this paper cites.
Self-referencing embedded strings (selfies): A 100% robust molecular string representation
Mario Krenn, Florian Häse, AkshatKumar Nigam, Pascal Friederich, and Alan Aspuru-Guzik. 2020 · 2020
Earlier work this paper cites.
Biobert: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020 · 2020
Earlier work this paper cites.
Behrt: transformer for electronic health records
Yikuan Li, Shishir Rao, José Roberto Ayala Solares, Abdelaali Hassaine, Rema Ramakrishnan, Dexter Canoy, Yajie Zhu, Kazem Rahimi, and Gholamreza Salimi-Khorshidi. 2020 · 2020
Earlier work this paper cites.
S2orc: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel S Weld. 2020 · 2020
Earlier work this paper cites.
Earthquake transformer—an attentive deep-learning model for simultaneous earthquake detection and phase picking
S Mostafa Mousavi, William L Ellsworth, Weiqiang Zhu, Lindsay Y Chuang, and Gregory C Beroza. 2020 · 2020
Earlier work this paper cites.
On the effectiveness of small, discriminatively pre-trained language representation models for biomedical text mining
Ibrahim Burak Ozyurt. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Biomegatron: larger biomedical domain language model
Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, and Raghav Mani. 2020 · 2020
Earlier work this paper cites.
Medicat: A dataset of medical images, captions, and textual references
Sanjay Subramanian, Lucy Lu Wang, Sachin Mehta, Ben Bogin, Madeleine van Zuylen, Sravanthi Parasa, Sameer Singh, Matt Gardner, and Hannaneh Hajishirzi. 2020 · 2020
Earlier work this paper cites.
Tabert: Pretraining for joint understanding of textual and tabular data
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020 · 2020
Earlier work this paper cites.
Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning
Jiaqi Chen, Jianheng Tang, Jinghui Qin, Xiaodan Liang, Lingbo Liu, Eric Xing, and Liang Lin. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Earlier work this paper cites.
Gakg: A multimodal geoscience academic knowledge graph
Cheng Deng, Yuting Jia, Hui Xu, Chong Zhang, Jingyao Tang, Luoyi Fu, Weinan Zhang, Haisong Zhang, Xinbing Wang, and Chenghu Zhou. 2021 · 2021
Earlier work this paper cites.
Text2mol: Cross-modal molecule retrieval with natural language queries
Carl Edwards, ChengXiang Zhai, and Heng Ji. 2021 · 2021
Earlier work this paper cites.
Prottrans: Toward understanding the language of life through self-supervised learning
Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, et al. 2021 · 2021
Earlier work this paper cites.
Capturing row and column semantics in transformer based question answering over tables
Michael Glass, Mustafa Canim, Alfio Gliozzo, Saneem Chemmengath, Vishwajeet Kumar, Rishav Chakravarti, Avirup Sil, Feifei Pan, Samarth Bharadwaj, and Nicolas Rodolfo Fauceglia. 2021 · 2021
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2021 · 2021
Earlier work this paper cites.
Learning the language of viral evolution and escape
Brian Hie, Ellen D Zhong, Bonnie Berger, and Bryan Bryson. 2021 · 2021
Earlier work this paper cites.
Extracting a knowledge base of mechanisms from covid-19 papers
Tom Hope, Aida Amini, David Wadden, Madeleine van Zuylen, Sravanthi Parasa, Eric Horvitz, Daniel S Weld, Roy Schwartz, and Hannaneh Hajishirzi. 2021 · 2021
Earlier work this paper cites.
Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition
Shih-Cheng Huang, Liyue Shen, Matthew P Lungren, and Serena Yeung. 2021 · 2021
Earlier work this paper cites.
Tabbie: Pretrained representations of tabular data
Hiroshi Iida, Dung Thai, Varun Manjunatha, and Mohit Iyyer. 2021 · 2021
Cited alongside, same era.
Dnabert: pre-trained bidirectional encoder representations from transformers model for dna-language in genome
Yanrong Ji, Zhihan Zhou, Han Liu, and Ramana V Davuluri. 2021 · 2021
Cited alongside, same era.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021 · 2021
Cited alongside, same era.
Mmbert: Multimodal bert pretraining for improved medical vqa
Yash Khare, Viraj Bagal, Minesh Mathew, Adithi Devi, U Deva Priyakumar, and CV Jawahar. 2021 · 2021
Cited alongside, same era.
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-chun Zhu. 2021 · 2021
Cited alongside, same era.
Large language models generate functional protein sequences across diverse families
Ali Madani, Ben Krause, Eric R Greene, Subu Subramanian, Benjamin P Mohr, James M Holton, Jose Luis Olmos, Caiming Xiong, Zachary Z Sun, Richard Socher, et al. 2023 · 2023
Later among the works it cites.
W-mae: Pre-trained weather model with masked autoencoder for multi-variable weather forecasting
Xin Man, Chenghong Zhang, Jin Feng, Changyu Li, and Jie Shao. 2023 · 2023
Later among the works it cites.
Med-flamingo: a multimodal medical few-shot learner
Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Eduardo Pontes Reis, and Pranav Rajpurkar. 2023 · 2023
Later among the works it cites.
Covid-twitter-bert: A natural language processing model to analyse covid-19 content on twitter
Martin Müller, Marcel Salathé, and Per E Kummervold. 2023 · 2023
Later among the works it cites.
Progen2: exploring the boundaries of protein language models
Erik Nijkamp, Jeffrey A Ruffolo, Eli N Weinstein, Nikhil Naik, and Ali Madani. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Explaining relationships between scientific documents
Kelvin Luu, Xinyi Wu, Rik Koncel-Kedziorski, Kyle Lo, Isabel Cachola, and Noah A Smith. 2021 · 2021
Cited alongside, same era.
Language models enable zero-shot prediction of the effects of mutations on protein function
Joshua Meier, Roshan Rao, Robert Verkuil, Jason Liu, Tom Sercu, and Alex Rives. 2021 · 2021
Cited alongside, same era.
Electramed: a new pre-trained language representation model for biomedical nlp
Giacomo Miolo, Giulio Mantoan, and Carlotta Orsenigo. 2021 · 2021
Cited alongside, same era.
Scifive: a text-to-text transformer model for biomedical literature
Long N Phan, James T Anibal, Hieu Tran, Shaurya Chanana, Erol Bahadroglu, Alec Peltekian, and Grégoire Altan-Bonnet. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Cited alongside, same era.
Msa transformer
Roshan M Rao, Jason Liu, Robert Verkuil, Joshua Meier, John Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives. 2021 · 2021
Cited alongside, same era.
Med-bert: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction
Laila Rasmy, Yang Xiang, Ziqian Xie, Cui Tao, and Degui Zhi. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
A symbolic characters aware model for solving geometry problems
Maizhen Ning, Qiu-Feng Wang, Kaizhu Huang, and Xiaowei Huang. 2023 · 2023
Later among the works it cites.
Catalyst energy prediction with catberta: Unveiling feature exploration strategies through large language models
Janghoon Ock, Chakradhar Guntuboina, and Amir Barati Farimani. 2023 · 2023
Later among the works it cites.
Biot5: Enriching cross-modal integration in biology with chemical knowledge and natural language associations
Qizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu, Kaiyuan Gao, Lijun Wu, Yingce Xia, and Rui Yan. 2023 · 2023
Later among the works it cites.
Xplainer: From x-ray observations to explainable zero-shot diagnosis
Chantal Pellegrini, Matthias Keicher, Ege Özsoy, Petra Jiraskova, Rickmer Braren, and Nassir Navab. 2023 · 2023
Later among the works it cites.
Bayesian optimization of catalysts with in-context learning
Mayk Caldas Ramos, Shane S Michtavy, Marc D Porosoff, and Andrew D White. 2023 · 2023
Later among the works it cites.
Andre Niyongabo Rubungo, Craig Arnold, Barry P Rand, and Adji Bousso Dieng. 2023 · 2023
Later among the works it cites.
Climatebert-netzero: Detecting and assessing net zero and reduction targets
Tobias Schimanski, Julia Bingler, Mathias Kraus, Camilla Hyslop, and Markus Leippold. 2023 · 2023
Later among the works it cites.
A general-purpose material property data extraction pipeline from large polymer corpora using natural language processing
Pranav Shetty, Arunkumar Chitteth Rajan, Chris Kuenneth, Sonakshi Gupta, Lakshmi Prerana Panchumarti, Lauren Holm, Chao Zhang, and Rampi Ramprasad. 2023 · 2023
Later among the works it cites.
Scirepeval: A multi-format benchmark for scientific document representations
Amanpreet Singh, Mike D’Arcy, Arman Cohan, Doug Downey, and Sergey Feldman. 2023 · 2023
Later among the works it cites.
Monte carlo thought search: Large language model querying for complex scientific reasoning in catalyst design
Henry Sprueill, Carl Edwards, Mariefel Olarte, Udishnu Sanyal, Heng Ji, and Sutanay Choudhury. 2023 · 2023
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, et al. 2023 · 2023
Later among the works it cites.
Interactive and explainable region-guided radiology report generation
Tim Tanida, Philip Müller, Georgios Kaissis, and Daniel Rueckert. 2023 · 2023
Later among the works it cites.
Transfer learning enables predictions in network biology
Christina V Theodoris, Ling Xiao, Anant Chopra, Mark D Chaffin, Zeina R Al Sayed, Matthew C Hill, Helene Mantineo, Elizabeth M Brydon, Zexian Zeng, X Shirley Liu, et al. 2023 · 2023
Later among the works it cites.
Chatclimate: Grounding conversational ai in climate science
Saeid Ashraf Vaghefi, Dominik Stammbach, Veruska Muccione, Julia Bingler, Jingwei Ni, Mathias Kraus, Simon Allen, Chiara Colesanti-Senni, Tobias Wekhof, Tobias Schimanski, et al. 2023 · 2023
Later among the works it cites.
Med-unic: Unifying cross-lingual medical vision-language pre-training by diminishing bias
Zhongwei Wan, Che Liu, Mi Zhang, Jie Fu, Benyou Wang, Sibo Cheng, Lei Ma, César Quilodrán-Casas, and Rossella Arcucci. 2023 · 2023
Later among the works it cites.
The future of chemistry is language
Andrew D White. 2023 · 2023
Later among the works it cites.
Towards generalist foundation model for radiology
Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023 · 2023
Later among the works it cites.
A systematic survey of chemical pre-trained models
Jun Xia, Yanqiao Zhu, Yuanqi Du, and Stan Z Li. 2023 · 2023
Later among the works it cites.
Darwin series: Domain specific large language models for natural science
Tong Xie, Yuwei Wan, Wei Huang, Zhenyu Yin, Yixuan Liu, Shaozhou Wang, Qingyuan Linghu, Chunyu Kit, Clara Grazian, Wenjie Zhang, et al. 2023 · 2023
Later among the works it cites.
Doctorglm: Fine-tuning your chinese doctor is not a herculean task
Honglin Xiong, Sheng Wang, Yitao Zhu, Zihao Zhao, Yuxiao Liu, Linlin Huang, Qian Wang, and Dinggang Shen. 2023 · 2023
Later among the works it cites.
Forge: pre-training open foundation models for science
Junqi Yin, Sajal Dash, Feiyi Wang, and Mallikarjun Shankar. 2023 · 2023
Later among the works it cites.
Cxr-clip: Toward large scale chest x-ray language-image pre-training
Kihyun You, Jawook Gu, Jiyeon Ham, Beomhee Park, Jiho Kim, Eun K Hong, Woonhyuk Baek, and Byungseok Roh. 2023 · 2023
Later among the works it cites.
Selformer: molecular representation learning via selfies language models
Atakan Yüksel, Erva Ulusoy, Atabey Ünlü, and Tunca Doğan. 2023 · 2023
Later among the works it cites.
Dnagpt: A generalized pretrained tool for multiple dna sequence analysis tasks
Daoan Zhang, Weitong Zhang, Bing He, Jianguo Zhang, Chenchen Qin, and Jianhua Yao. 2023a · 2023
Later among the works it cites.
Uni-mol: A universal 3d molecular representation learning framework
Gengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng, Hongteng Xu, Zhewei Wei, Linfeng Zhang, and Guolin Ke. 2023 · 2023
Later among the works it cites.
Genslms: Genome-scale language models reveal sars-cov-2 evolutionary dynamics
Maxim Zvyagin, Alexander Brace, Kyle Hippe, Yuntian Deng, Bin Zhang, Cindy Orozco Bohorquez, Austin Clyde, Bharat Kale, Danilo Perez-Rivera, Heng Ma, et al. 2023 · 2023
Later among the works it cites.
Prot2text: Multimodal protein’s function generation with gnns and transformers
Hadi Abdine, Michail Chatzianastasis, Costas Bouyioukos, and Michalis Vazirgiannis. 2024 · 2024
Closest in time.
Hippocrates: An open-source framework for advancing large language models in healthcare
Emre Can Acikgoz, Osman Batur İnce, Rayene Bench, Arda Anıl Boz, İlker Kesen, Aykut Erdem, and Erkut Erdem. 2024 · 2024
Closest in time.
Meta-designing quantum experiments with language models
Sören Arlt, Haonan Duan, Felix Li, Sang Michael Xie, Yuhuai Wu, and Mario Krenn. 2024 · 2024
Closest in time.
Llemma: An open language model for mathematics
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q Jiang, Jia Deng, Stella Biderman, and Sean Welleck. 2024 · 2024
Closest in time.
Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, and Sung Ju Hwang. 2024 · 2024
Closest in time.
M3d: Advancing 3d medical image analysis with multi-modal large language models
Fan Bai, Yuxin Du, Tiejun Huang, Max Q-H Meng, and Bo Zhao. 2024 · 2024
Closest in time.
Biomedlm: A 2.7 b parameter language model trained on biomedical text
Elliot Bolton, Abhinav Venigalla, Michihiro Yasunaga, David Hall, Betty Xiong, Tony Lee, Roxana Daneshjou, Jonathan Frankle, Percy Liang, Michael Carbin, et al. 2024 · 2024
Closest in time.
Augmenting large language models with chemistry tools
Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. 2024 · 2024
Closest in time.
Transforming the bootstrap: Using transformers to compute scattering amplitudes in planar n= 4 super yang-mills theory
Tianji Cai, Garrett W Merz, François Charton, Niklas Nolte, Matthias Wilhelm, Kyle Cranmer, and Lance J Dixon. 2024 · 2024
Closest in time.
Bidirectional generation of structure and properties through a single molecular foundation model
Jinho Chang and Jong Chul Ye. 2024 · 2024
Closest in time.
Self-supervised learning on millions of primary rna sequences from 72 vertebrates improves sequence-based rna splicing prediction
Ken Chen, Yue Zhou, Maolin Ding, Yu Wang, Zhixiang Ren, and Yuedong Yang. 2024 · 2024
Closest in time.
Bartsmiles: Generative masked language models for molecular representations
Gayane Chilingaryan, Hovhannes Tamoyan, Ani Tevosyan, Nelly Babayan, Lusine Khondkaryan, Karen Hambardzumyan, Zaven Navoyan, Hrant Khachatrian, and Armen Aghajanyan. 2024 · 2024
Closest in time.
A 5’ utr language model for decoding untranslated regions of mrna and function predictions
Yanyi Chu, Dan Yu, Yupeng Li, Kaixuan Huang, Yue Shen, Le Cong, Jason Zhang, and Mengdi Wang. 2024 · 2024
Closest in time.
scgpt: toward building a foundation model for single-cell multi-omics using generative ai
Haotian Cui, Chloe Wang, Hassaan Maan, Kuan Pang, Fengning Luo, Nan Duan, and Bo Wang. 2024 · 2024
Closest in time.
Marg: Multi-agent review generation for scientific papers
Mike D’Arcy, Tom Hope, Larry Birnbaum, and Doug Downey. 2024 · 2024
Closest in time.
K2: A foundation language model for geoscience knowledge understanding and utilization
Cheng Deng, Tianhang Zhang, Zhongmou He, Qiyuan Chen, Yuanyuan Shi, Yi Xu, Luoyi Fu, Weinan Zhang, Xinbing Wang, Chenghu Zhou, et al. 2024 · 2024
Closest in time.
David Fitzek, Yi Hong Teoh, Hin Pok Fung, Gebremedhin A Dagnew, Ejaaz Merali, M Schuyler Moss, Benjamin MacLellan, and Roger G Melko. 2024 · 2024
Closest in time.
Shantanu Ghosh, Clare B Poynton, Shyam Visweswaran, and Kayhan Batmanghelich. 2024 · 2024
Closest in time.
Tora: A tool-integrated reasoning agent for mathematical problem solving
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen. 2024 · 2024
Closest in time.
Building astrobert, a language model for astronomy & astrophysics
F Grezes, S Blanco-Cuaresma, A Accomazzi, MJ Kurtz, G Shapurian, E Henneken, CS Grant, DM Thompson, R Chyla, S McDonald, et al. 2024 · 2024
Closest in time.
Fine-tuned language models generate stable inorganic materials as text
Nate Gruver, Anuroop Sriram, Andrea Madotto, Andrew Gordon Wilson, C Lawrence Zitnick, and Zachary Ulissi. 2024 · 2024
Closest in time.
Xuemei Gu and Mario Krenn. 2024 · 2024
Closest in time.
Large-scale foundation model on single-cell transcriptomics
Minsheng Hao, Jing Gong, Xin Zeng, Chiming Liu, Yucheng Guo, Xingyi Cheng, Taifeng Wang, Jianzhu Ma, Xuegong Zhang, and Le Song. 2024 · 2024
Closest in time.
Physbert: A text embedding model for physics scientific literature
Thorsten Hellert, João Montenegro, and Andrea Pollastro. 2024 · 2024
Closest in time.
A survey of pre-trained language models for processing scientific text
Xanh Ho, Anh Khoa Duong Nguyen, An Tuan Dao, Junfeng Jiang, Yuki Chida, Kaito Sugimoto, Huy Quoc To, Florian Boudin, and Akiko Aizawa. 2024 · 2024
Closest in time.
Leveraging large language models for predictive chemistry
Kevin Maik Jablonka, Philippe Schwaller, Andres Ortega-Guerrero, and Berend Smit. 2024 · 2024
Closest in time.
Graph chain-of-thought: Augmenting large language models by reasoning on graphs
Bowen Jin, Chulin Xie, Jiawei Zhang, Kashob Kumar Roy, Yu Zhang, Zheng Li, Ruirui Li, Xianfeng Tang, Suhang Wang, Yu Meng, and Jiawei Han. 2024 · 2024
Closest in time.
Transparent medical image ai via an image–text foundation model grounded in medical literature
Chanwoo Kim, Soham U Gadgil, Alex J DeGrave, Jesutofunmi A Omiye, Zhuo Ran Cai, Roxana Daneshjou, and Su-In Lee. 2024 · 2024
Closest in time.
Biomistral: A collection of open-source pretrained large language models for medical domains
Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024 · 2024
Closest in time.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. 2024 · 2024
Closest in time.
Taiyi: a bilingual fine-tuned large language model for diverse biomedical tasks
Ling Luo, Jinzhong Ning, Yingwen Zhao, Zhijun Wang, Zeyuan Ding, Peng Chen, Weiru Fu, Qinyu Han, Guangtao Xu, Yunzhi Qiu, et al. 2024 · 2024
Closest in time.
Prollama: A protein large language model for multi-task protein language processing
Liuzhenghao Lv, Zongying Lin, Hao Li, Yuyang Liu, Jiaxi Cui, Calvin Yu-Chian Chen, Li Yuan, and Yonghong Tian. 2024 · 2024
Closest in time.
Relative molecule self-attention transformer
Łukasz Maziarka, Dawid Majchrowski, Tomasz Danel, Piotr Gaiński, Jacek Tabor, Igor Podolak, Paweł Morkisz, and Stanisław Jastrzębski. 2024 · 2024
Closest in time.
Are large language models superhuman chemists?
Adrian Mirza, Nawaf Alampara, Sreekanth Kunchapu, Benedict Emoekabu, Aswanth Krishnan, Mara Wilhelmi, Macjonathan Okereke, Juliane Eberhardt, Amir Mohammad Elahi, Maximilian Greiner, et al. 2024 · 2024
Closest in time.
Leveraging biomolecule and natural language through multi-modal learning: A survey
Qizhi Pei, Lijun Wu, Kaiyuan Gao, Jinhua Zhu, Yue Wang, Zun Wang, Tao Qin, and Rui Yan. 2024 · 2024
Closest in time.
Astrollama-chat: Scaling astrollama with conversational and diverse datasets
Ernest Perkowski, Rui Pan, Tuan Dung Nguyen, Yuan-Sen Ting, Sandor Kruk, Tong Zhang, Charlie O’Neill, Maja Jablonska, Zechang Sun, Michael J Smith, et al. 2024 · 2024
Closest in time.
Bimedix: Bilingual medical mixture of experts llm
Sara Pieri, Sahal Shaji Mullappilly, Fahad Shahbaz Khan, Rao Muhammad Anwer, Salman Khan, Timothy Baldwin, and Hisham Cholakkal. 2024 · 2024
Closest in time.
Mathematical discoveries from program search with large language models
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al. 2024 · 2024
Closest in time.
Capabilities of gemini models in medicine
Khaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn, Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, et al. 2024 · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, YK Li, Y Wu, and Daya Guo. 2024 · 2024
Closest in time.
Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers
Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. 2024 · 2024
Closest in time.
Shiven Sinha, Ameya Prabhu, Ponnurangam Kumaraguru, Siddharth Bhat, and Matthias Bethge. 2024 · 2024
Closest in time.
Chemreasoner: Heuristic search over a large language model’s knowledge space using quantum-chemical feedback
Henry W Sprueill, Carl Edwards, Khushbu Agarwal, Mariefel V Olarte, Udishnu Sanyal, Conrad Johnston, Hongbin Liu, Heng Ji, and Sutanay Choudhury. 2024 · 2024
Closest in time.
Bioclip: A vision foundation model for the tree of life
Samuel Stevens, Jiaman Wu, Matthew J Thompson, Elizabeth G Campolongo, Chan Hee Song, David Edward Carlyn, Li Dong, Wasila M Dahdul, Charles Stewart, Tanya Berger-Wolf, et al. 2024 · 2024
Closest in time.
Saprot: protein language modeling with structure-aware vocabulary
Jin Su, Chenchen Han, Yuyang Zhou, Junjie Shan, Xibin Zhou, and Fajie Yuan. 2024 · 2024
Closest in time.
Xraygpt: Chest radiographs summarization using medical vision-language models
Omkar Thawkar, Abdelrahman Shaker, Sahal Shaji Mullappilly, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, Jorma Laaksonen, and Fahad Shahbaz Khan. 2024 · 2024
Closest in time.
Openmathinstruct-1: A 1.8 million math instruction tuning dataset
Shubham Toshniwal, Ivan Moshkov, Sean Narenthiran, Daria Gitman, Fei Jia, and Igor Gitman. 2024 · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong. 2024 · 2024
Closest in time.
Towards generalist biomedical ai
Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann, Mohamed Amin, Pi-Chuan Chang, Andrew Carroll, Charles Lau, Ryutaro Tanno, Ira Ktena, et al. 2024 · 2024
Closest in time.
Cellplm: Pre-training of cell language model beyond single cells
Hongzhi Wen, Wenzhuo Tang, Xinnan Dai, Jiayuan Ding, Wei Jin, Yuying Xie, and Jiliang Tang. 2024 · 2024
Closest in time.
Pmc-llama: toward building open-source language models for medicine
Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang, Weidi Xie, and Yanfeng Wang. 2024 · 2024
Closest in time.
Me llama: Foundation large language models for medical applications
Qianqian Xie, Qingyu Chen, Aokun Chen, Cheng Peng, Yan Hu, Fongci Lin, Xueqing Peng, Jimin Huang, Jeffrey Zhang, Vipina Keloth, et al. 2024 · 2024
Closest in time.
Benchmarking retrieval-augmented generation for medicine
Guangzhi Xiong, Qiao Jin, Zhiyong Lu, and Aidong Zhang. 2024 · 2024
Closest in time.
Bmretriever: Tuning large language models as better biomedical text retrievers
Ran Xu, Wenqi Shi, Yue Yu, Yuchen Zhuang, Yanqiao Zhu, May D Wang, Joyce C Ho, Chao Zhang, and Carl Yang. 2024 · 2024
Closest in time.
Urbanclip: Learning text-enhanced urban region profiling with contrastive language-image pretraining from the web
Yibo Yan, Haomin Wen, Siru Zhong, Wei Chen, Haodong Chen, Qingsong Wen, Roger Zimmermann, and Yuxuan Liang. 2024 · 2024
Closest in time.
Internlm-math: Open math large language models toward verifiable reasoning
Huaiyuan Ying, Shuo Zhang, Linyang Li, Zhejian Zhou, Yunfan Shao, Zhaoye Fei, Yichuan Ma, Jiawei Hong, Kuikun Liu, Ziyi Wang, et al. 2024 · 2024
Closest in time.
Chemdfm: Dialogue foundation model for chemistry
Zihan Zhao, Da Ma, Lu Chen, Liangtai Sun, Zihao Li, Hongshen Xu, Zichen Zhu, Su Zhu, Shuai Fan, Guodong Shen, et al. 2024 · 2024
Closest in time.
Can large language models transform computational social science?
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024 · 2024
Closest in time.
Automated chemical reaction extraction from scientific literature
Jiang Guo, A Santiago Ibanez-Lopez, Hanyu Gao, Victor Quach, Connor W Coley, Klavs F Jensen, and Regina Barzilay. 2022 · 2045
Closest in time.
The era5 global reanalysis
Hans Hersbach, Bill Bell, Paul Berrisford, Shoji Hirahara, András Horányi, Joaquín Muñoz-Sabater, Julien Nicolas, Carole Peubey, Raluca Radu, Dinand Schepers, et al. 2020 · 2049
Closest in time.
Molgpt: molecular generation using a transformer-decoder model
Viraj Bagal, Rishal Aggarwal, PK Vinod, and U Deva Priyakumar. 2022 · 2076
Closest in time.