Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable generalizability, such as understanding arbitrary entities and relations.
Genia corpus—a semantically annotated corpus for bio-textmining
J-D Kim, Tomoko Ohta, Yuka Tateisi, and Jun’ichi Tsujii · 2003
Earlier work this paper cites.
Introduction to the bio-entity recognition task at jnlpba
Nigel Collier and Jin-Dong Kim · 2004
Earlier work this paper cites.
Ace 2004 multilingual training corpus
Alexis Mitchell, Stephanie Strassel, Shudong Huang, and Ramez Zakhary · 2005
Earlier work this paper cites.
Ace 2005 multilingual training corpus
Christopher Walker, Stephanie Strassel, Julie Medero, and Kazuaki Maeda · 2006
Earlier work this paper cites.
Evaluating the state-of-the-art in automatic de-identification
Özlem Uzuner, Yuan Luo, and Peter Szolovits · 2007
Earlier work this paper cites.
Open information extraction from the web
Oren Etzioni, Michele Banko, Stephen Soderland, and Daniel S Weld · 2008
Earlier work this paper cites.
Overview of biocreative ii gene mention recognition
Larry Smith, Lorraine K Tanabe, Cheng-Ju Kuo, I Chung, Chun-Nan Hsu, Yu-Shi Lin, Roman Klinger, Christoph M Friedrich, Kuzman Ganchev, Manabu Torii, et al · 2008
Earlier work this paper cites.
2010 i2b2/va challenge on concepts, assertions, and relations in clinical text
Özlem Uzuner, Brett R South, Shuying Shen, and Scott L DuVall · 2011
Earlier work this paper cites.
Asgard: A portable architecture for multilingual dialogue systems
Jingjing Liu, Panupong Pasupat, Scott Cyphers, and Jim Glass · 2013
Earlier work this paper cites.
Semeval-2013 task 9: Extraction of drug-drug interactions from biomedical texts (ddiextraction 2013)
Isabel Segura-Bedmar, Paloma Martínez Fernández, and María Herrero Zazo · 2013
Earlier work this paper cites.
Evaluating temporal relations in clinical text: 2012 i2b2 challenge
Weiyi Sun, Anna Rumshisky, and Ozlem Uzuner · 2013
Earlier work this paper cites.
Ontonotes release 5.0 ldc2013t19
Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, et al · 2013
Earlier work this paper cites.
Ncbi disease corpus: a resource for disease name recognition and concept normalization
Rezarta Islamaj Doğan, Robert Leaman, and Zhiyong Lu · 2014
Earlier work this paper cites.
Task 2: Share/clef ehealth evaluation lab 2014
Danielle L Mowery, Sumithra Velupillai, Brett R South, Lee Christensen, David Martinez, Liadh Kelly, Lorraine Goeuriot, Noemie Elhadad, Sameer Pradhan, Guergana Savova, et al · 2014
Earlier work this paper cites.
Anatomical entity mention recognition at literature scale
Sampo Pyysalo and Sophia Ananiadou · 2014
Earlier work this paper cites.
Polyglot-ner: Massive multilingual named entity recognition
Rami Al-Rfou, Vivek Kulkarni, Bryan Perozzi, and Steven Skiena · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
The chemdner corpus of chemicals and drugs and its annotation principles
Martin Krallinger, Obdulia Rabal, Florian Leitner, Miguel Vazquez, David Salgado, Zhiyong Lu, Robert Leaman, Yanan Lu, Donghong Ji, Daniel M Lowe, et al · 2015
Earlier work this paper cites.
Automated systems for the de-identification of longitudinal clinical narratives: Overview of 2014 i2b2/uthealth shared task track 1
Amber Stubbs, Christopher Kotfila, and Özlem Uzuner · 2015
Earlier work this paper cites.
Broad Twitter corpus: A diverse named entity recognition resource
Leon Derczynski, Kalina Bontcheva, and Ian Roberts · 2016
Earlier work this paper cites.
An annotated corpus for machine reading of instructions in wet lab protocols
Chaitanya Kulkarni, Wei Xu, Alan Ritter, and Raghu Machiraju · 2016
Earlier work this paper cites.
Biocreative v cdr task corpus: a resource for chemical disease relation extraction
Jiao Li, Yueping Sun, Robin J Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J Mattingly, Thomas C Wiegers, and Zhiyong Lu · 2016
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji · 2017
Earlier work this paper cites.
Clinical named entity recognition using deep learning models
Yonghui Wu, Min Jiang, Jun Xu, Degui Zhi, and Hua Xu · 2017
Cited alongside, same era.
Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction
Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi · 2018
Cited alongside, same era.
A corpus with multi-level annotations of patients, interventions and outcomes to support language processing for medical literature
Benjamin Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang, Iain J Marshall, Ani Nenkova, and Byron C Wallace · 2018
Cited alongside, same era.
Crossweigh: Training named entity tagger from imperfect annotations
Zihan Wang, Jingbo Shang, Liyuan Liu, Lihao Lu, Jiacheng Liu, and Jiawei Han · 2019
Cited alongside, same era.
The sofc-exp corpus and neural approaches to information extraction in the materials science domain, 2020
Annemarie Friedrich, Heike Adel, Federico Tomazic, Johannes Hingerl, Renou Benteau, Anika Maruscyk, and Lukas Lange · 2020
Cited alongside, same era.
Biored: a rich biomedical relation extraction dataset
Ling Luo, Po-Ting Lai, Chih-Hsuan Wei, Cecilia N Arighi, and Zhiyong Lu · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, et al · 2022
Later among the works it cites.
Language models in the loop: Incorporating prompting into weak supervision
Ryan Smith, Jason A Fries, Braden Hancock, and Stephen H Bach · 2022
Later among the works it cites.
MultiNERD: A multilingual, multi-genre and fine-grained dataset for named entity recognition (and disambiguation)
Simone Tedeschi and Roberto Navigli · 2022
Later among the works it cites.
Named Entity Recognition in Twitter: A Dataset and Analysis on Short-Term Temporal Shifts
Asahi Ushio, Leonardo Neves, Vitor Silva, Francesco. Barbieri, and Jose Camacho-Collados · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Cited alongside, same era.
2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records
Sam Henry, Kevin Buchan, Michele Filannino, Amber Stubbs, and Ozlem Uzuner · 2020
Cited alongside, same era.
SciREX: A challenge dataset for document-level information extraction
Sarthak Jain, Madeleine van Zuylen, Hannaneh Hajishirzi, and Iz Beltagy · 2020
Cited alongside, same era.
Named entity recognition and relation detection for biomedical information extraction
Nadeesha Perera, Matthias Dehmer, and Frank Emmert-Streib · 2020
Cited alongside, same era.
Code and named entity recognition in StackOverflow
Jeniya Tabassum, Mounica Maddela, Wei Xu, and Alan Ritter · 2020
Cited alongside, same era.
Few-NERD: A few-shot named entity recognition dataset
Ning Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, and Zhiyuan Liu · 2021
Cited alongside, same era.
Crossner: Evaluating cross-domain named entity recognition
Zihan Liu, Yan Xu, Tiezheng Yu, Wenliang Dai, Ziwei Ji, Samuel Cahyawijaya, Andrea Madotto, and Pascale Fung · 2021
Cited alongside, same era.
Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, and Xudong Shen · 2022
Later among the works it cites.
Tasteset–recipe dataset and food entities recognition benchmark
Ania Wróblewska, Agnieszka Kaliska, Maciej Pawłowski, Dawid Wiśniewski, Witold Sosnowski, and Agnieszka Ławrynowicz · 2022
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Closest in time.
Distilling large language models for biomedical knowledge extraction: A case study on adverse drug events, 2023
Yu Gu, Sheng Zhang, Naoto Usuyama, Yonas Woldesenbet, Cliff Wong, Praneeth Sanapathi, Mu Wei, Naveen Valluri, Erika Strandberg, Tristan Naumann, and Hoifung Poon · 2023
Closest in time.
Runwei Guan, Ka Lok Man, Feifan Chen, Shanliang Yao, Rongsheng Hu, Xiaohui Zhu, Jeremy Smith, Eng Gee Lim, and Yutao Yue · 2023
Closest in time.
The false promise of imitating proprietary llms, 2023
Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song · 2023
Closest in time.
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Cheng-Yu Hsieh, Chun-Liang Li, Chih-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister · 2023
Closest in time.
Impossible distillation: from low-quality model to high-quality dataset & model for summarization and paraphrasing, 2023
Jaehun Jung, Peter West, Liwei Jiang, Faeze Brahman, Ximing Lu, Jillian Fisher, Taylor Sorensen, and Yejin Choi · 2023
Closest in time.
Bo Li, Gexiang Fang, Yang Yang, Quansen Wang, Wei Ye, Wen Zhao, and Shikun Zhang · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Closest in time.
Finer: Financial named entity recognition dataset and weak-supervision model
Agam Shah, Ruchit Vithani, Abhinav Gullapalli, and Sudheer Chava · 2023
Closest in time.
Pushing the limits of chatgpt on nlp tasks, 2023
Xiaofei Sun, Linfeng Dong, Xiaoya Li, Zhen Wan, Shuhe Wang, Tianwei Zhang, Jiwei Li, Fei Cheng, Lingjuan Lyu, Fei Wu, and Guoyin Wang · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Closest in time.
Zero-shot information extraction via chatting with chatgpt
Xiang Wei, Xingyu Cui, Ning Cheng, Xiaobin Wang, Xin Zhang, Shen Huang, Pengjun Xie, Jinan Xu, Yufeng Chen, Meishan Zhang, et al · 2023
Closest in time.
A comprehensive capability analysis of gpt-3 and gpt-3.5 series models
Junjie Ye, Xuanting Chen, Nuo Xu, Can Zu, Zekai Shao, Shichun Liu, Yuhan Cui, Zeyang Zhou, Chao Gong, Yang Shen, et al · 2023
Closest in time.