Fetching the paper…
Reading the bibliography…
With the rapid development of NLP research, leaderboards have emerged as one tool to track the performance of various systems on various NLP tasks.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019 · 1905
Earlier work this paper cites.
Bootstrap methods: another look at the jackknife
Bradley Efron. 1992 · 1992
Earlier work this paper cites.
Stacked generalizations: When does it work?
Kai Ming Ting and Ian H. Witten. 1997 · 1997
Earlier work this paper cites.
Thumbs up? sentiment classification using machine learning techniques
Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. 2002 · 2002
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
Kristina Toutanova, Dan Klein, Christopher D. Manning, and Yoram Singer. 2003 · 2003
Earlier work this paper cites.
Xglue: A new benchmark dataset for cross-lingual pre-training, understanding and generation
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, et al. 2020 · 2004
Earlier work this paper cites.
A high-performance semi-supervised learning method for text chunking
Rie Ando and Tong Zhang. 2005 · 2005
Earlier work this paper cites.
Ian Tenney, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang, Mahima Pushkarna, Carey Radebaugh, Emily Reif, et al. 2020 · 2008
Earlier work this paper cites.
Generalized minimum bayes risk system combination
Kevin Duh, Katsuhito Sudoh, Xianchao Wu, Hajime Tsukada, and Masaaki Nagata. 2011 · 2011
Earlier work this paper cites.
Minimum Bayes-risk system combination
Jesús González-Rubio, Alfons Juan, and Francisco Casacuberta. 2011 · 2011
Earlier work this paper cites.
Towards automatic error analysis of machine translation output
Maja Popović and Hermann Ney. 2011 · 2011
Earlier work this paper cites.
Flert: Document-level features for named entity recognition
Stefan Schweter and Alan Akbik. 2020 · 2011
Earlier work this paper cites.
Blast: A tool for error analysis of machine translation output
Sara Stymne. 2011 · 2011
Earlier work this paper cites.
Baselines and bigrams: Simple, good sentiment and topic classification
Sida Wang and Christopher Manning. 2012 · 2012
Earlier work this paper cites.
Semantic parsing on Freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013 · 2013
Earlier work this paper cites.
Identifying intention posts in discussion forums
Zhiyuan Chen, Bing Liu, Meichun Hsu, Malu Castellanos, and Riddhiman Ghosh. 2013 · 2013
Earlier work this paper cites.
Explaining the stars: Weighted multiple-instance learning for aspect-based sentiment analysis
Nikolaos Pappas and Andrei Popescu-Belis. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Long short-term memory neural networks for Chinese word segmentation
Xinchi Chen, Xipeng Qiu, Chenxi Zhu, Pengfei Liu, and Xuanjing Huang. 2015 · 2015
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016 · 2016
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
A closer look at data bias in neural extractive summarization models
Ming Zhong, Danqing Wang, Pengfei Liu, Xipeng Qiu, and Xuanjing Huang. 2019 · 2019
Later among the works it cites.
Findings of the 2020 conference on machine translation (WMT20)
Loïc Barrault, Magdalena Biesialska, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, Tom Kocmi, Philipp Koehn, Chi-kiu Lo, Nikola Ljubešić, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Santanu Pal, Matt Post, and Marcos Zampieri. 2020 · 2020
Later among the works it cites.
Re-evaluating evaluation in text summarization
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020 · 2020
Later among the works it cites.
Utility is in the eye of the user: A critique of NLP leaderboards
Kawin Ethayarajh and Dan Jurafsky. 2020 · 2020
Later among the works it cites.
Interpretable multi-dataset evaluation for named entity recognition
Jinlan Fu, Pengfei Liu, and Graham Neubig. 2020a · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017 · 2017
Cited alongside, same era.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018 · 2018
Cited alongside, same era.
Pooled contextualized embeddings for named entity recognition
Alan Akbik, Tanja Bergmann, and Roland Vollgraf. 2019 · 2019
Cited alongside, same era.
Rethinking generalization of neural models: A named entity recognition case study
Jinlan Fu, Pengfei Liu, and Qi Zhang. 2020b · 2020
Later among the works it cites.
RethinkCWS: Is Chinese word segmentation a solved task?
Jinlan Fu, Pengfei Liu, Qi Zhang, and Xuanjing Huang. 2020c · 2020
Later among the works it cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Later among the works it cites.
LUKE: Deep contextualized entity representations with entity-aware self-attention
Ikuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda, and Yuji Matsumoto. 2020 · 2020
Later among the works it cites.
Does syntax matter? a strong baseline for aspect-based sentiment analysis with roberta
Junqi Dai, Hang Yan, Tianxiang Sun, Pengfei Liu, and Xipeng Qiu. 2021 · 2021
Closest in time.
Spanner: Named entity re-/recognition as span prediction
Jinlan Fu, Xuanjing Huang, and Pengfei Liu. 2021 · 2021
Closest in time.
The gem benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh D Dhole, et al. 2021 · 2021
Closest in time.
Genie: A leaderboard for human-in-the-loop evaluation of text generation
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A Smith, and Daniel S Weld. 2021 · 2021
Closest in time.
Simcls: A simple framework for contrastive learning of abstractive summarization
Yixin Liu and Pengfei Liu. 2021 · 2021
Closest in time.
Refsum: Refactoring neural summarization
Yixin Liu, Dou Ziyi, and Pengfei Liu. 2021 · 2021
Closest in time.
Towards more fine-grained and reliable NLP performance prediction
Zihuiwen Ye, Pengfei Liu, Jinlan Fu, and Graham Neubig. 2021 · 2021
Closest in time.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Closest in time.