Fetching the paper…
Reading the bibliography…
Large language models (LLMs), typically designed as a function of next-word prediction, have excelled across extensive NLP tasks.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 1910
Earlier work this paper cites.
The use of the area under the roc curve in the evaluation of machine learning algorithms
Andrew P Bradley · 1997
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia · 2001
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin · 2003
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F. Brown, Stephen A. Della Pietra, Vincent J. Della Pietra, and Robert L. Mercer · 2003
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Evaluation of compositional distributional semantic models on full sentences through semantic relatedness and textual entailment
SemEval-2014 · 2014
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli · 2014
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D. Manning · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng · 2016
Earlier work this paper cites.
Variations of the similarity function of textrank for automated summarization
Federico Barrios, Federico López, Luis Argerich, and Rosa Wachenchauzer · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
NewsQA: A machine comprehension dataset
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Earlier work this paper cites.
Mind the GAP: A balanced corpus of gendered ambiguous pronouns
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge · 2018
Earlier work this paper cites.
Marian: Fast neural machine translation in C++
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Tomasz Dwojak, Hieu Hoang, Kenneth Heafield, Tom Neckermann, Frank Seide, Ulrich Germann, Alham Fikri Aji, Nikolay Bogoychev, André F. T. Martins, and Alexandra Birch · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Earlier work this paper cites.
Faithful to the original: Fact aware neural abstractive summarization
Ziqiang Cao, Furu Wei, Wenjie Li, and Sujian Li · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Wikihow: A large scale text summarization dataset
Mahnaz Koupaee and William Yang Wang · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
PAWS: Paraphrase adversaries from word scrambling
Yuan Zhang, Jason Baldridge, and Luheng He · 2019
Cited alongside, same era.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Cited alongside, same era.
DREAM: A challenge data set and models for dialogue-based reading comprehension
Kai Sun, Dian Yu, Jianshu Chen, Dong Yu, Yejin Choi, and Claire Cardie · 2019
Cited alongside, same era.
Looking beyond sentence-level natural language inference for question answering and text summarization
Anshuman Mishra, Dhruvesh Patel, Aparna Vijayakumar, Xiang Lorraine Li, Pavan Kapanipathi, and Kartik Talamadupula · 2021
Later among the works it cites.
Muppet: Massive multi-task representations with pre-finetuning
Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta · 2021
Later among the works it cites.
SimCSE: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen · 2021
Later among the works it cites.
Get your vitamin C! robust fact verification with contrastive evidence
Tal Schuster, Adam Fisch, and Regina Barzilay · 2021
Later among the works it cites.
Compression, transduction, and creation: A unified framework for evaluating natural language generation
Mingkai Deng, Bowen Tan, Zhengzhong Liu, Eric Xing, and Zhiting Hu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark · 2019
Cited alongside, same era.
Neural text summarization: A critical evaluation
Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher · 2019
Cited alongside, same era.
A simple recipe towards reducing hallucination in neural surface realisation
Feng Nie, Jin-Ge Yao, Jinpeng Wang, Rong Pan, and Chin-Yew Lin · 2019
Cited alongside, same era.
SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization
Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer · 2019
Cited alongside, same era.
Read + verify: Machine reading comprehension with unanswerable questions
Minghao Hu, Furu Wei, Yuxing Peng, Zhen Huang, Nan Yang, and Dongsheng Li · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Cited alongside, same era.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Cited alongside, same era.
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev · 2021
Later among the works it cites.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu · 2021
Later among the works it cites.
Do we know what we don’t know? studying unanswerable questions beyond SQuAD 2.0
Elior Sulem, Jamaal Hay, and Dan Roth · 2021
Later among the works it cites.
DocNLI: A large-scale dataset for document-level natural language inference
Wenpeng Yin, Dragomir Radev, and Caiming Xiong · 2021
Later among the works it cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant · 2021
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus · 2022
Later among the works it cites.
MVP: multi-task supervised pre-training for natural language generation
Tianyi Tang, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei · 2022
Later among the works it cites.
Toward a ’Standard Model’ of Machine Learning
Zhiting Hu and Eric P. Xing · 2022
Later among the works it cites.
SummaC: Re-visiting NLI-based models for inconsistency detection in summarization
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst · 2022
Later among the works it cites.
SMART: sentences as basic units for text evaluation
Reinald Kim Amplayo, Peter J. Liu, Yao Zhao, and Shashi Narayan · 2022
Later among the works it cites.
TRUE: Re-evaluating factual consistency evaluation
Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias · 2022
Later among the works it cites.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han · 2022
Later among the works it cites.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong · 2022
Later among the works it cites.
Chatgpt survey: Performance on NLP datasets, Mar 2023
Matúš Pikuliak · 2023
Closest in time.
AlignScore: Evaluating factual consistency with a unified alignment function
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu · 2023
Closest in time.
Cappy: Outperforming and boosting large multi-task lms with a small scorer
Bowen Tan, Yun Zhu, Lijuan Liu, Eric Xing, Zhiting Hu, and Jindong Chen · 2023
Closest in time.
The flan collection: Designing data and methods for effective instruction tuning
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts · 2023
Closest in time.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Closest in time.
G-eval: NLG evaluation using GPT-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu · 2023
Closest in time.
Human-like summarization evaluation with chatgpt
Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, and Xiaojun Wan · 2023
Closest in time.