The 2020 bilingual, bi-directional webnlg+ shared task overview and evaluation results (webnlg+ 2020)
Thiago Castro Ferreira, Claire Gardent, Nikolai Ilinykh, Chris van der Lee, Simon Mille, Diego Moussallem, and Anastasia Shimorina · 2020
Later among the works it cites.
Pre-training tasks for embedding-based large-scale retrieval
Original
Wei-Cheng Chang, Felix X Yu, Yin-Wen Chang, Yiming Yang, and Sanjiv Kumar · 2020
Later among the works it cites.
GoEmotions: A dataset of fine-grained emotions
Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi · 2020
Later among the works it cites.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace · 2020
Later among the works it cites.
Wikilingua: A new benchmark dataset for multilingual abstractive summarization
Claire Cardie Faisal Ladhak, Esin Durmus and Kathleen McKeown · 2020
Later among the works it cites.
Benchmarking meaning representations in neural semantic parsing
Jiaqi Guo, Qian Liu, Jian-Guang Lou, Zhenwen Li, Xueqing Liu, Tao Xie, and Ting Liu · 2020
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith · 2020
Later among the works it cites.
XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson · 2020
Later among the works it cites.
Neural CRF model for sentence alignment in text simplification
Chao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong, and Wei Xu · 2020
Later among the works it cites.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy · 2020
Later among the works it cites.
Scaling laws for neural language models
Original
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
UNIFIEDQA: Crossing format boundaries with a single QA system
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi · 2020
Later among the works it cites.
COGS: A compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen · 2020
Later among the works it cites.
Task-specific objectives of pre-trained language models for dialogue adaptation, 2020
Junlong Li, Zhuosheng Zhang, Hai Zhao, Xi Zhou, and Xiang Zhou · 2020
Later among the works it cites.
CommonGen: A constrained text generation challenge for generative commonsense reasoning
Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren · 2020
Later among the works it cites.
Dialoglue: A natural language understanding benchmark for task-oriented dialogue, 2020
Shikib Mehri, Mihail Eric, and Dilek Hakkani-Tur · 2020
Later among the works it cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2020
Later among the works it cites.
Document ranking with a pretrained sequence-to-sequence model
Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin · 2020
Later among the works it cites.
Totto: A controlled table-to-text generation dataset
Original
Ankur P Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das · 2020
Later among the works it cites.
Dart: Open-domain structured data record to text generation
Original
Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Nazneen Fatema Rajani, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Murori Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, and Richard Socher · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Later among the works it cites.
Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset, 2020
Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan · 2020
Later among the works it cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer · 2020
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale
Original
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2020
Later among the works it cites.
Glu variants improve transformer, 2020
Noam Shazeer · 2020
Later among the works it cites.
Exploring and predicting transferability across NLP tasks
Tu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, and Mohit Iyyer · 2020
Later among the works it cites.
Gradient surgery for multi-task learning, 2020
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Later among the works it cites.
Muppet: Massive multi-task representations with pre-finetuning, 2021
Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta · 2021
Closest in time.
Efficiently identifying task groupings for multi-task learning
Original
Christopher Fifty, Ehsan Amid, Zhe Zhao, Tianhe Yu, Rohan Anil, and Chelsea Finn · 2021
Closest in time.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou · 2021
Closest in time.
What’s in your head? emergent behaviour in multi-task transformer models
Mor Geva, Uri Katz, Aviv Ben-Arie, and Jonathan Berant · 2021
Closest in time.
Unlocking compositional generalization in pre-trained models using intermediate representations, 2021
Jonathan Herzig, Peter Shaw, Ming-Wei Chang, Kelvin Guu, Panupong Pasupat, and Yuan Zhang · 2021
Closest in time.
Question answering infused pre-training of general-purpose contextualized representations, 2021
Robin Jia, Mike Lewis, and Luke Zettlemoyer · 2021
Closest in time.
Quiz-Style Question Generation for News Stories , pp. 2501–2511
Adam D. Lelkes, Vinh Q. Tran, and Cong Yu · 2021
Closest in time.
Unicorn on rainbow: A universal commonsense reasoning model on a new multitask benchmark
Nicholas Lourie, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2021
Closest in time.
StylePTB: A compositional benchmark for fine-grained controllable text style transfer
Yiwei Lyu, Paul Pu Liang, Hai Pham, Eduard Hovy, Barnabás Póczos, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2021
Closest in time.
AgreeSum: Agreement-oriented multi-document summarization
Richard Yuanzhe Pang, Adam Lelkes, Vinh Tran, and Cong Yu · 2021
Closest in time.
KILT: a benchmark for knowledge intensive language tasks
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel · 2021
Closest in time.
Few-shot question answering by pretraining span selection
Ori Ram, Yuval Kirstain, Jonathan Berant, Amir Globerson, and Omer Levy · 2021
Closest in time.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2021
Closest in time.
Get your vitamin C! robust fact verification with contrastive evidence
Tal Schuster, Adam Fisch, and Regina Barzilay · 2021
Closest in time.
Process for adapting language models to society (palms) with values-targeted datasets
Original
Irene Solaiman and Christy Dennison · 2021
Closest in time.
Ranking Transfer Languages with Pragmatically-Motivated Features for Multilingual Sentiment Analysis
Jimin Sun, Hwijeen Ahn, Chan Young Park, Yulia Tsvetkov, and David R. Mortensen · 2021
Closest in time.
Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models
Zirui Wang, Yulia Tsvetkov, Orhan Firat, and Yuan Cao · 2021
Closest in time.
Finetuned language models are zero-shot learners, 2021
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le · 2021
Closest in time.
DocNLI: A large-scale dataset for document-level natural language inference
Wenpeng Yin, Dragomir Radev, and Caiming Xiong · 2021
Closest in time.
Scaling vision transformers
Original
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2021
Closest in time.
When do you need billions of words of pretraining data?
Yian Zhang, Alex Warstadt, Xiaocheng Li, and Samuel R. Bowman · 2021
Closest in time.
Identifying beneficial task relations for multi-task learning in deep neural networks
Joachim Bingel and Anders Søgaard · 2026
Closest in time.