Fetching the paper…
Reading the bibliography…
Training and evaluating language models increasingly requires the construction of meta-datasets --diverse collections of curated data with clear provenance.
Developing a test collection for biomedical word sense disambiguation
M Weeber, J G Mork, and A R Aronson · 2001
Earlier work this paper cites.
Mining MEDLINE: abstracts, sentences, or phrases?
J Ding, D Berleant, D Nettleton, and E Wurtele · 2002
Earlier work this paper cites.
The genia corpus: An annotated research abstract corpus in molecular biology domain
Tomoko Ohta, Yuka Tateisi, and Jin-Dong Kim · 2002
Earlier work this paper cites.
A multi-layered, xml-based approach to the integration of linguistic and semantic annotations
Paul Buitelaar, Thierry Declerck, Bogdan Sacaleanu, Špela Vintar, Diana Raileanu, and Claudia Crispi · 2003
Earlier work this paper cites.
Introduction to the bio-entity recognition task at JNLPBA
Nigel Collier and Jin-Dong Kim · 2004
Earlier work this paper cites.
Overview of biocreative: critical assessment of information extraction for biology, 2005
Lynette Hirschman, Alexander Yeh, Christian Blaschke, and Alfonso Valencia · 2005
Earlier work this paper cites.
GENETAG: a tagged corpus for gene/protein named entity recognition
Lorraine Tanabe, Natalie Xie, Lynne H Thom, Wayne Matten, and W John Wilbur · 2005
Earlier work this paper cites.
Mutationfinder: a high-performance system for extracting point mutation mentions from text
J. Gregory Caporaso, William A Baumgartner, David A Randolph, K. Bretonnel Cohen, and Lawrence Hunter · 2007
Earlier work this paper cites.
Relex–relation extraction using dependency parse trees
Katrin Fundel, Robert Küffner, and Ralf Zimmer · 2007
Earlier work this paper cites.
Results of the fifth edition of the BioASQ challenge
Anastasios Nentidis, Konstantinos Bougiatiotis, Anastasia Krithara, Georgios Paliouras, and Ioannis Kakadiaris · 2007
Earlier work this paper cites.
Measures of semantic similarity and relatedness in the biomedical domain
Ted Pedersen, Serguei VS Pakhomov, Siddharth Patwardhan, and Christopher G Chute · 2007
Earlier work this paper cites.
Bioinfer: a corpus for information extraction in the biomedical domain
Sampo Pyysalo, Filip Ginter, Juho Heimonen, Jari Bj"orne, Jorma Boberg, Jouni J"arvinen, and Tapio Salakoski · 2007
Earlier work this paper cites.
Evaluating the state-of-the-art in automatic de-identification
Özlem Uzuner, Yuan Luo, and Peter Szolovits · 2007
Earlier work this paper cites.
Osirisv1.2: a named entity recognition system for sequence variants of genes in biomedical literature
Laura I Furlong, Holger Dach, Martin Hofmann-Apitius, and Ferran Sanz · 2008
Earlier work this paper cites.
Chemical names: Terminological resources and corpora annotation
Corinna Kol’arik, Roman Klinger, Christoph M Friedrich, Martin Hofmann-Apitius, and Juliane Fluck · 2008
Earlier work this paper cites.
Identifying patient smoking status from medical discharge records
Ozlem Uzuner, Ira Goldstein, Yuan Luo, and Isaac Kohane · 2008
Earlier work this paper cites.
The bioscope corpus: biomedical texts annotated for uncertainty, negation and their scopes
Veronika Vincze, Gy"orgy Szarvas, Rich’ard Farkas, Gy"orgy M’ora, and J’anos Csirik · 2008
Earlier work this paper cites.
Data fusion
Jens Bleiholder and Felix Naumann · 2009
Earlier work this paper cites.
Overview of bionlp’09 shared task on event extraction
Jin-Dong Kim, Tomoko Ohta, Sampo Pyysalo, Yoshinobu Kano, and Jun’ichi Tsujii · 2009
Earlier work this paper cites.
Overview of BioNLP’09 shared task on event extraction
Jin-Dong Kim, Tomoko Ohta, Sampo Pyysalo, Yoshinobu Kano, and Jun’ichi Tsujii · 2009
Earlier work this paper cites.
Static relations: a piece in the biomedical information extraction puzzle
Sampo Pyysalo, Tomoko Ohta, Jin-Dong Kim, and Jun’ichi Tsujii · 2009
Earlier work this paper cites.
Recognizing obesity and comorbidities in sparse data
Ozlem Uzuner · 2009
Earlier work this paper cites.
Linnaeus: a species name identification system for biomedical literature
Martin Gerner, Goran Nenadic, and Casey M Bergman · 2010
Earlier work this paper cites.
An empirical evaluation of resources for the identification of diseases and adverse effects in biomedical literature
Harsha Gurulingappa, Roman Klinger, Martin Hofmann-Apitius, and Juliane Fluck · 2010
Earlier work this paper cites.
Event extraction for post-translational modifications
Tomoko Ohta, Sampo Pyysalo, Makoto Miwa, Jin-Dong Kim, and Jun’ichi Tsujii · 2010
Earlier work this paper cites.
Semantic similarity and relatedness between clinical terms: an experimental study
Serguei Pakhomov, Bridget McInnes, Terrence Adam, Ying Liu, Ted Pedersen, and Genevieve B Melton · 2010
Earlier work this paper cites.
Extracting medication information from clinical text
Ozlem Uzuner, Imre Solti, and Eithon Cadag · 2010
Earlier work this paper cites.
Exploiting mesh indexing in medline to generate a data set for word sense disambiguation
Antonio J Jimeno-Yepes, Bridget T McInnes, and Alan R Aronson · 2011
Earlier work this paper cites.
Overview of genia event task in bionlp shared task 2011
Jin-Dong Kim, Yue Wang, Toshihisa Takagi, and Akinori Yonezawa · 2011
Earlier work this paper cites.
Overview of the epigenetics and post-translational modifications (EPI) task of BioNLP shared task 2011
Tomoko Ohta, Sampo Pyysalo, and Jun’ichi Tsujii · 2011
Earlier work this paper cites.
Overview of the infectious diseases (ID) task of BioNLP shared task 2011
Sampo Pyysalo, Tomoko Ohta, Rafal Rak, Dan Sullivan, Chunhong Mao, Chunxia Wang, Bruno Sobral, Jun’ichi Tsujii, and Sophia Ananiadou · 2011
Earlier work this paper cites.
Overview of the entity relations (rel) supporting task of bionlp shared task 2011
Sampo Pyysalo, Tomoko Ohta, and Jun’ichi Tsujii · 2011
Earlier work this paper cites.
Challenges in the association of human single nucleotide polymorphism mentions with unique database identifiers
Philippe Thomas, Roman Klinger, Laura Furlong, Martin Hofmann-Apitius, and Christoph Friedrich · 2011
Earlier work this paper cites.
2010 i2b2/va challenge on concepts, assertions, and relations in clinical text
Ozlem Uzuner, Brett R. South, Shuying Shen, and Scott L. DuVall · 2011
Earlier work this paper cites.
Brat: a web-based tool for nlp-assisted text annotation
Pontus Stenetorp, Sampo Pyysalo, Goran Topić, Tomoko Ohta, Sophia Ananiadou, and Jun’ichi Tsujii · 2012
Earlier work this paper cites.
Annotating and evaluating text for stem cell research
Mariana Neves, Alexander Damaschun, Andreas Kurtz, and Ulf Leser · 2012
Earlier work this paper cites.
Open-domain anatomical entity mention detection
Tomoko Ohta, Sampo Pyysalo, Jun’ichi Tsujii, and Sophia Ananiadou · 2012
Earlier work this paper cites.
Event extraction across multiple levels of biological organization
Sampo Pyysalo, Tomoko Ohta, Makoto Miwa, Han-Cheol Cho, Jun’ichi Tsujii, and Sophia Ananiadou · 2012
Earlier work this paper cites.
Evaluating the state of the art in coreference resolution for electronic medical records
Ozlem Uzuner, Andreea Bodnari, Shuying Shen, Tyler Forbush, John Pestian, and Brett R South · 2012
Earlier work this paper cites.
The eu-adr corpus: Annotated drugs, diseases, targets, and their relationships
Erik M. van Mulligen, Annie Fourrier-Reglat, David Gurwitz, Mariam Molokhia, Ainhoa Nieto, Gianluca Trifiro, Jan A. Kors, and Laura I. Furlong · 2012
Earlier work this paper cites.
Redundancy in electronic health record corpora: analysis, impact on text mining performance and mitigation strategies
Raphael Cohen, Michael Elhadad, and Noémie Elhadad · 2013
Earlier work this paper cites.
Bioc: a minimalist approach to interoperability for biomedical text processing
Donald C Comeau, Rezarta Islamaj Doğan, Paolo Ciccarese, Kevin Bretonnel Cohen, Martin Krallinger, Florian Leitner, Zhiyong Lu, Yifan Peng, Fabio Rinaldi, Manabu Torii, et al · 2013
Earlier work this paper cites.
The ddi corpus: An annotated corpus with pharmacological substances and drug–drug interactions
María Herrero-Zazo, Isabel Segura-Bedmar, Paloma Martínez, and Thierry Declerck · 2013
Earlier work this paper cites.
The Genia event extraction shared task, 2013 edition - overview
Jin-Dong Kim, Yue Wang, and Yamamoto Yasunori · 2013
Earlier work this paper cites.
GRO task: Populating the gene regulation ontology with events and relations
Jung-jae Kim, Xu Han, Vivian Lee, and Dietrich Rebholz-Schuhmann · 2013
Earlier work this paper cites.
Overview of the pathway curation (PC) task of BioNLP shared task 2013
Tomoko Ohta, Sampo Pyysalo, Rafal Rak, Andrew Rowley, Hong-Woo Chun, Sung-Jae Jung, Sung-Pil Choi, Sophia Ananiadou, and Jun’ichi Tsujii · 2013
Earlier work this paper cites.
Overview of the cancer genetics (CG) task of BioNLP shared task 2013
Sampo Pyysalo, Tomoko Ohta, and Sophia Ananiadou · 2013
Earlier work this paper cites.
Annotating the biomedical literature for the human variome
Karin Verspoor, Antonio Jimeno Yepes, Lawrence Cavedon, Tara McIntosh, Asha Herten-Crabb, Zo"e Thomas, and John-Paul Plazzer · 2013
Earlier work this paper cites.
tmvar: a text mining approach for extracting sequence variants in biomedical literature
Chih-Hsuan Wei, Bethany R Harris, Hung-Yu Kao, and Zhiyong Lu · 2013
Earlier work this paper cites.
Detecting mirna mentions and relations in biomedical literature
Shweta Bagewadi, Tamara Bobi’c, Martin Hofmann-Apitius, Juliane Fluck, and Roman Klinger · 2014
Earlier work this paper cites.
Ncbi disease corpus: A resource for disease name recognition and concept normalization
Rezarta Islamaj Dogan, Robert Leaman, and Zhiyong Lu · 2014
Earlier work this paper cites.
Discourse complements lexical semantics for non-factoid answer reranking
Peter Jansen, Mihai Surdeanu, and Peter Clark · 2014
Earlier work this paper cites.
Creation of a new longitudinal corpus of clinical narratives
Vishesh Kumar, Amber Stubbs, Stanley Shaw, and Özlem Uzuner · 2014
Earlier work this paper cites.
The QUAERO French medical corpus: A ressource for medical entity recognition and normalization
Aurélie Névéol, Cyril Grouin, Jeremy Leixa, Sophie Rosset, and Pierre Zweigenbaum · 2014
Earlier work this paper cites.
Anatomical entity mention recognition at literature scale
Sampo Pyysalo and Sophia Ananiadou · 2014
Earlier work this paper cites.
Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research
Àlex Bravo, Janet Piñero, N’uria Queralt-Rosinach, Michael Rautschka, and Laura I Furlong · 2015
Cited alongside, same era.
Cadec: A corpus of adverse drug event annotations
Sarvnaz Karimi, Alejandro Metke-Jimenez, Madonna Kemp, and Chen Wang · 2015
Cited alongside, same era.
A multilingual gold-standard corpus for biomedical concept recognition: the Mantra GSC
Jan A Kors, Simon Clematide, Saber A Akhondi, Erik M van Mulligen, and Dietrich Rebholz-Schuhmann · 2015
Cited alongside, same era.
The chemdner corpus of chemicals and drugs and its annotation principles
Martin Krallinger, Obdulia Rabal, Florian Leitner, Miguel Vazquez, David Salgado, Zhiyong Lu, Robert Leaman, Yanan Lu, Donghong Ji, Daniel M. Lowe, Roger A. Sayle, Riza Theresa Batista-Navarro, Rafal Rak, Torsten Huber, Tim Rockt"aschel, S’ergio Matos, David Campos, Buzhou Tang, Hua Xu, Tsendsuren Munkhdalai, Keun Ho Ryu, S. V. Ramanan, Senthil Nathan, Slavko Zitnik, Marko Bajec, Lutz Weber, Matthias Irmer, Saber A. Akhondi, Jan A. Kors, Shuo Xu, Xin An, Utpal Kumar Sikdar, Asif Ekbal, Masaharu Yoshioka, Thaer M. Dieb, Miji Choi, Karin Verspoor, Madian Khabsa, C. Lee Giles, Hongfang Liu, Komandur Elayavilli Ravikumar, Andre Lamurias, Francisco M. Couto, Hong-Jie Dai, Richard Tzong-Han Tsai, Caglar Ata, Tolga Can, Anabel Usi’e, Rui Alves, Isabel Segura-Bedmar, Paloma Mart’inez, Julen Oyarzabal, and Alfonso Valencia · 2015
2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records
Sam Henry, Kevin Buchan, Michele Filannino, Amber Stubbs, and Ozlem Uzuner · 2020
Later among the works it cites.
Explainable automated fact-checking for public health claims
Neema Kotonya and Francesca Toni · 2020
Later among the works it cites.
Effective transfer learning for identifying similar questions: Matching user questions to covid-19 faqs
Rabal-O. Lourenço A. Krallinger, M · 2020
Later among the works it cites.
Chia, a large annotated corpus of clinical trial eligibility criteria
Fabr’ıcio Kury, Alex Butler, Chi Yuan, Li-heng Fu, Yingcheng Sun, Hao Liu, Ida Sim, Simona Carini, and Chunhua Weng · 2020
Later among the works it cites.
Multi-xscience: A large-scale dataset for extreme multi-document summarization of scientific articles, 2020
Yao Lu, Yue Dong, and Laurent Charlin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Automated systems for the de-identification of longitudinal clinical narratives: Overview of 2014 i2b2/uthealth shared task track 1
Amber Stubbs, Christopher Kotfila, and Özlem Uzuner · 2015
Cited alongside, same era.
An overview of the bioasq large-scale biomedical semantic indexing and question answering competition
George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, et al · 2015
Cited alongside, same era.
GNormPlus: An integrative approach for tagging genes, gene families, and protein domains
Chih-Hsuan Wei, Hung-Yu Kao, and Zhiyong Lu · 2015
Cited alongside, same era.
Named entity recognition in swedish medical journals with deep bidirectional character-based lstms
Simon Almgren, Sean Pavlov, and Olof Mogren · 2016
Cited alongside, same era.
Automatic semantic classification of scientific literature according to the hallmarks of cancer
Simon Baker, Ilona Silins, Yufan Guo, Imran Ali, Johan H"ogberg, Ulla Stenius, and Anna Korhonen · 2016
Cited alongside, same era.
Normalising medical concepts in social media texts by learning semantic representation
Nut Limsopatham and Nigel Collier · 2016
Cited alongside, same era.
Seth detects and normalizes genetic variants in text
Philippe Thomas, Tim Rockt"aschel, J"org Hakenberg, Yvonne Lichtblau, and Ulf Leser · 2016
Cited alongside, same era.
Biomedical big data: New models of control over access, use and governance
Effy Vayena and Alessandro Blasimme · 2017
Cited alongside, same era.
Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results
Antonio Miranda-Escalada, Eulàlia Farré, and Martin Krallinger · 2020
Later among the works it cites.
Overview of automatic clinical coding: Annotations, guidelines, and solutions for non-english clinical cases at codiesp track of clef ehealth 2020
Antonio Miranda-Escalada, Aitor Gonzalez-Agirre, Jordi Armengol-Estapé, and Martin Krallinger · 2020
Later among the works it cites.
BioMRC: A dataset for biomedical machine reading comprehension
Dimitris Pappas, Petros Stavropoulos, Ion Androutsopoulos, and Ryan McDonald · 2020
Later among the works it cites.
Biomedical concept relatedness – a large EHR-based benchmark
Claudia Schulz, Josh Levy-Kramer, Camille Van Assel, Miklos Kepes, and Nils Hammerla · 2020
Later among the works it cites.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi · 2020
Later among the works it cites.
Comprehensive named entity recognition on CORD-19 with distant or weak supervision
Xuan Wang, Xiangchen Song, Yingjun Guan, Bangzheng Li, and Jiawei Han · 2020
Later among the works it cites.
Effective crowd-annotation of participants, interventions, and outcomes in the text of clinical trial reports
Markus Zlabinger, Marta Sabou, Sebastian Hofst"atter, and Allan Hanbury · 2020
Later among the works it cites.
Muppet: Massive multi-task representations with pre-finetuning
Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta · 2021
Later among the works it cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow
Sid Black, Gao Leo, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Later among the works it cites.
Memorization vs. generalization : Quantifying data leakage in NLP performance evaluation
Aparna Elangovan, Jiayuan He, and Karin Verspoor · 2021
Later among the works it cites.
A framework for few-shot language model evaluation, September 2021
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2021
Later among the works it cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf · 2021
Later among the works it cites.
Template-free prompt tuning for few-shot ner
Ruotian Ma, Xin Zhou, Tao Gui, Yiding Tan, Qi Zhang, and Xuanjing Huang · 2021
Later among the works it cites.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna · 2021
Later among the works it cites.
Scifive: a text-to-text transformer model for biomedical literature
Long N Phan, James T Anibal, Hieu Tran, Shaurya Chanana, Erol Bahadroglu, Alec Peltekian, and Grégoire Altan-Bonnet · 2021
Later among the works it cites.
Changing the world by changing the data
Anna Rogers · 2021
Later among the works it cites.
“everyone wants to do the model work, not the data work”: Data cascades in high-stakes ai
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo · 2021
Later among the works it cites.
Fine-tuning large neural language models for biomedical natural language processing
Robert Tinn, Hao Cheng, Yu Gu, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon · 2021
Later among the works it cites.
Massive choice, ample tasks (MaChAmp): A toolkit for multi-task learning in NLP
Rob van der Goot, Ahmet Üstün, Alan Ramponi, Ibrahim Sharaf, and Barbara Plank · 2021
Later among the works it cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki · 2021
Later among the works it cites.
Hunflair: an easy-to-use tool for state-of-the-art biomedical named entity recognition
Leon Weber, Mario Sänger, Jannes Münchmeyer, Maryam Habibi, Ulf Leser, and Alan Akbik · 2021
Later among the works it cites.
Ku-dmis at bioasq 9: Data-centric and model-centric approaches for biomedical question answering
Wonjin Yoon, Jaehyo Yoo, Sumin Seo, Mujeen Sung, Minbyul Jeong, Gangwoo Kim, and Jaewoo Kang · 2021
Later among the works it cites.
A clinical trials corpus annotated with UMLS entities to enhance the access to evidence-based medicine
Leonardo Campillos-Llanos, Ana Valverde-Mateos, Adri’an Capllonch-Carri’on, and Antonio Moreno-Sandoval · 2021
Later among the works it cites.
Overview of the biocreative vii litcovid track: multi-label topic classification for covid-19 literature annotation
Qingyu Chen, Alexis Allot, Robert Leaman, Rezarta Islamaj Doğan, and Zhiyong Lu · 2021
Later among the works it cites.
Overview of bioasq 2021-mesinesp track. evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials
Luis Gasco, Anastasios Nentidis, Anastasia Krithara, Darryl Estrada-Zavala, Renato Toshiyuki Murasaki, Elena Primo-Peña, Cristina Bojo Canales, Georgios Paliouras, Martin Krallinger, et al · 2021
Later among the works it cites.
Nlm-chem, a new resource for chemical entity recognition in pubmed full text literature
Rezarta Islamaj, Robert Leaman, Sun Kim, Dongseop Kwon, Chih-Hsuan Wei, Donald C Comeau, Yifan Peng, David Cissel, Cathleen Coss, Carol Fisher, et al · 2021
Later among the works it cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits · 2021
Later among the works it cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2021
Later among the works it cites.
Paramed: a parallel corpus for english–chinese translation in the biomedical domain
Boxiang Liu and Liang Huang · 2021
Later among the works it cites.
Covid-19 named entity recognition for vietnamese
Thinh Hung Truong, Mai Hoang Dao, and Dat Quoc Nguyen · 2021
Later among the works it cites.
Large language models are zero-shot clinical information extractors
Monica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim, and David Sontag · 2022
Closest in time.
PromptSource: An integrated development environment and repository for natural language prompts
Stephen H. Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-David, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Alan Fries, Maged S. Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Dragomir Radev, Mike Tian-Jian Jiang, and Alexander M. Rush · 2022
Closest in time.
Dataset debt in biomedical language modeling
Jason Fries, Natasha Seelam, Gabriel Altay, Leon Weber, Myungsun Kang, Debajyoti Datta, Ruisi Su, Samuele Garda, Bo Wang, Simon Ott, Matthias Samwald, and Wojciech Kusa · 2022
Closest in time.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon · 2022
Closest in time.
Data governance in the age of large-scale data-driven language technology
Yacine Jernite, Huu Nguyen, Stella Biderman, Anna Rogers, Maraim Masoud, Valentin Danchev, Samson Tan, Alexandra Sasha Luccioni, Nishant Subramani, Gérard Dupont, Jesse Dodge, Kyle Lo, Zeerak Talat, Isaac Johnson, Dragomir Radev, Somaieh Nikpoor, Jörg Frohberg, Aaron Gokaslan, Peter Henderson, Rishi Bommasani, and Margaret Mitchell · 2022
Closest in time.
In-boxbart: Get instructions into biomedical multi-task learning
Mihir Parmar, Swaroop Mishra, Mirali Purohit, Man Luo, M Hassan Murad, and Chitta Baral · 2022
Closest in time.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush · 2022
Closest in time.
Benchmarking generalization via in-context instructions on 1,600+ language tasks
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al · 2022
Closest in time.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le · 2022
Closest in time.
Linkbert: Pretraining language models with document links
Michihiro Yasunaga, Jure Leskovec, and Percy Liang · 2022
Closest in time.
CBLUE: A Chinese biomedical language understanding evaluation benchmark
Ningyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li, Xin Shang, Kangping Yin, Chuanqi Tan, Jian Xu, Fei Huang, Luo Si, Yuan Ni, Guotong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan, Linfeng Li, Jun Yan, Hongying Zan, Kunli Zhang, Buzhou Tang, and Qingcai Chen · 2022
Closest in time.
DisTEMIST corpus: detection and normalization of disease mentions in spanish clinical cases, April 2022
Luis Gasco, Eulàlia Farré, Antonio Miranda-Escalada, Salvador Lima, and Martin Krallinger · 2022
Closest in time.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon · 2022
Closest in time.
Biored: A comprehensive biomedical relation extraction dataset
Ling Luo, Po-Ting Lai, Chih-Hsuan Wei, Cecilia N Arighi, and Zhiyong Lu · 2022
Closest in time.
Biored: A comprehensive biomedical relation extraction dataset
Ling Luo, Po-Ting Lai, Chih-Hsuan Wei, Cecilia N. Arighi, and Zhiyong Lu · 2022
Closest in time.
In-boxbart: Get instructions into biomedical multi-task learning
Mihir Parmar, Swaroop Mishra, Mirali Purohit, Man Luo, M Hassan Murad, and Chitta Baral · 2022
Closest in time.
tmvar 3.0: an improved variant concept recognition and normalization tool, 2022
Chih-Hsuan Wei, Alexis Allot, Kevin Riehle, Aleksandar Milosavljevic, and Zhiyong Lu · 2022
Closest in time.
Pmc-patients: A large-scale dataset of patient notes and relations extracted from case reports in pubmed central, 2022
Zhengyun Zhao, Qiao Jin, and Sheng Yu · 2022
Closest in time.