Fetching the paper…
Reading the bibliography…
Existing question answering (QA) datasets derived from electronic health records (EHR) are artificially generated and consequently fail to capture realistic physician information needs.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Scispacy: Fast and robust models for biomedical natural language processing
Mark Neumann, Daniel King, Iz Beltagy, and Waleed Ammar. 2019 · 1902
Earlier work this paper cites.
Let’s ask again: Refine network for automatic question generation
Preksha Nema, Akash Kumar Mohankumar, Mitesh M. Khapra, Balaji Vasan Srinivasan, and Balaraman Ravindran. 2019 · 1909
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, T. J. Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeff Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
An evaluation of information-seeking behaviors of general pediatricians
Donna D’Alessandro, Clarence Kreiter, and Michael Peterson. 2004 · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
Development, implementation, and a cognitive evaluation of a definitional question answering system for physicians
Hong Yu, Minsuk Lee, David R. Kaufman, John W. Ely, Jerome A. Osheroff, George Hripcsak, and James J. Cimino. 2007 · 2007
Earlier work this paper cites.
What can natural language processing do for clinical decision support?
Dina Demner-Fushman, Wendy Chapman, and Clement Mcdonald. 2009 · 2009
Earlier work this paper cites.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze. 2021 · 2009
Earlier work this paper cites.
Re-evaluating evaluation in text summarization
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Peng fei Liu, and Graham Neubig. 2020 · 2010
Earlier work this paper cites.
Good question! statistical ranking for question generation
Michael Heilman and Noah A. Smith. 2010 · 2010
Earlier work this paper cites.
Askhermes: An online question answering system for complex clinical questions
Yonggang Cao, F. Liu, Pippa M Simpson, Lamont D. Antieau, Andrew S. Bennett, James J. Cimino, John W. Ely, and Hong Yu. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Earlier work this paper cites.
Comparing automatic evaluation measures for image description
Desmond Elliott and Frank Keller. 2014 · 2014
Earlier work this paper cites.
Linguistic considerations in automatic question generation
Karen Mazidi and Rodney D. Nielsen. 2014 · 2014
Earlier work this paper cites.
Means: A medical question-answering system combining nlp techniques and semantic web technologies
Asma Ben Abacha and Pierre Zweigenbaum. 2015 · 2015
Earlier work this paper cites.
Towards topic-to-question generation
Yllias Chali and Sadid A. Hasan. 2015 · 2015
Earlier work this paper cites.
Deep questions without deep understanding
Igor Labutov, Sumit Basu, and Lucy Vanderwende. 2015 · 2015
Earlier work this paper cites.
Neural responding machine for short-text conversation
Lifeng Shang, Zhengdong Lu, and Hang Li. 2015 · 2015
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, H Lehman Li-wei, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016 · 2016
Cited alongside, same era.
A corpus and cloze evaluation for deeper understanding of commonsense stories
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016 · 2016
Cited alongside, same era.
Multiresolution recurrent neural networks: An application to dialogue response generation
Iulian Vlad Serban, Tim Klinger, Gerald Tesauro, Kartik Talamadupula, Bowen Zhou, Yoshua Bengio, and Aaron C. Courville. 2016 · 2016
Cited alongside, same era.
Overview of the medical question answering task at trec 2017 liveqa
Asma Ben Abacha, Eugene Agichtein, Yuval Pinter, and Dina Demner-Fushman. 2017 · 2017
Cited alongside, same era.
Learning to ask: Neural question generation for reading comprehension
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Semantic graphs for generating deep questions
Liangming Pan, Yuxi Xie, Yansong Feng, Tat-Seng Chua, and Min-Yen Kan. 2020 · 2020
Later among the works it cites.
Training question answering models from synthetic data
Raul Puri, Ryan Spring, Mohammad Shoeybi, Mostofa Patwary, and Bryan Catanzaro. 2020 · 2020
Later among the works it cites.
ProphetNet: Predicting future n-gram for sequence-to-SequencePre-training
Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xinya Du, Junru Shao, and Claire Cardie. 2017 · 2017
Cited alongside, same era.
Question generation for question answering
Nan Duan, Duyu Tang, Peng Chen, and Ming Zhou. 2017 · 2017
Cited alongside, same era.
Why we need new evaluation metrics for NLG
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Question answering and question generation as dual tasks
Duyu Tang, Nan Duan, Tao Qin, and Ming Zhou. 2017 · 2017
Cited alongside, same era.
emrqa: A large corpus for question answering on electronic medical records
Anusri Pampari, Preethi Raghavan, Jennifer Liang, and Jian Peng. 2018 · 2018
Cited alongside, same era.
From eliza to xiaoice: Challenges and opportunities with social chatbots
Heung-Yeung Shum, Xiaodong He, and Di Li. 2018 · 2018
Cited alongside, same era.
Neural models for key phrase extraction and question generation
Sandeep Subramanian, Tong Wang, Xingdi Yuan, Saizheng Zhang, Adam Trischler, and Yoshua Bengio. 2018 · 2018
Cited alongside, same era.
Clicr: a dataset of clinical case reports for machine reading comprehension
Simon Suster and Walter Daelemans. 2018 · 2018
Cited alongside, same era.
Question-driven summarization of answers to consumer health questions
Max E. Savery, Asma Ben Abacha, Soumya Gayen, and Dina Demner-Fushman. 2020 · 2020
Later among the works it cites.
Clinical reading comprehension: A thorough analysis of the emrQA dataset
Xiang Yue, Bernal Jimenez Gutierrez, and Huan Sun. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Question answering with long multiple-span answers
Ming Zhu, Aman Ahuja, Da-Cheng Juan, Wei Wei, and Chandan K. Reddy. 2020 · 2020
Later among the works it cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Xiaodong Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2021 · 2021
Later among the works it cites.
Entity guided question generation with contextual structure and sequence information capturing
Qingbao Huang, Mingyi Fu, Linzhang Mo, Yi Cai, Jingyun Xu, Pijian Li, Qing Li, and Ho-fung Leung. 2021 · 2021
Later among the works it cites.
What would it take to get biomedical QA systems into practice?
Gregory Kell, Iain Marshall, Byron Wallace, and Andre Jaun. 2021 · 2021
Later among the works it cites.
Does bert pretrained on clinical notes reveal sensitive data?
Eric P. Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, and Byron C. Wallace. 2021 · 2021
Later among the works it cites.
Quiz-style question generation for news stories
Adam D. Lelkes, Vinh Q. Tran, and Cong Yu. 2021 · 2021
Later among the works it cites.
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021 · 2021
Later among the works it cites.
Cooperative learning of zero-shot machine reading comprehension
Hongyin Luo, Seunghak Yu, Shang-Wen Li, and James R. Glass. 2021 · 2021
Later among the works it cites.
Mixqg: Neural question generation with mixed answer types
Lidiya Murakhovs’ka, Chien Sheng Wu, Tong Niu, Wenhao Liu, and Caiming Xiong. 2021 · 2021
Later among the works it cites.
emrKBQA: A clinical knowledge-base question answering dataset
Preethi Raghavan, Jennifer J Liang, Diwakar Mahajan, Rachita Chandra, and Peter Szolovits. 2021 · 2021
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2021 · 2021
Later among the works it cites.
Cliniqg4qa: Generating diverse questions for domain adaptation of clinical question answering
Xiang Yue, Xinliang Frederick Zhang, Ziyu Yao, Simon Lin, and Huan Sun. 2021 · 2021
Later among the works it cites.
Promptsource: An integrated development environment and repository for natural language prompts
Stephen H. Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-David, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Alan Fries, Maged S. Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Dragomir Radev, Mike Tian-Jian Jiang, and Alexander M. Rush. 2022 · 2022
Closest in time.
Towards generalizable methods for automating risk score calculation
Jennifer J. Liang, Eric Lehman, Ananya S. Iyengar, Diwakar Mahajan, Preethi Raghavan, Cindy Y. Chang, and Peter Szolovits. 2022 · 2022
Closest in time.