Transformer-xl: Attentive language models beyond a fixed-length context
Original
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
Original
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. 2019 · 1904
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Original
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Megatron-LM: Training multi-billion parameter language models using gpu model parallelism
Original
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Original
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 1910
Earlier work this paper cites.
CCNet: Extracting high quality monolingual datasets from web crawl data
Original
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2019 · 1911
Earlier work this paper cites.
ABOUT ML: Annotation and benchmarking on understanding and transparency of machine learning lifecycles
Original
Inioluwa Deborah Raji and Jingying Yang. 2019 · 1912
Earlier work this paper cites.
Sony corp. of america v. universal city studios, inc
1984 · 1984
Earlier work this paper cites.
Scaling laws for neural language models
Original
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
NLTK: The Natural Language Toolkit
Edward Loper and Steven Bird. 2002 · 2002
Earlier work this paper cites.
Generation of random permutations of given number of elements using random sampling numbers
C. Radhakrishna Rao. 1961 · 2002
Earlier work this paper cites.
Kelly v. arriba soft corp
2003 · 2003
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan. 2003 · 2003
Earlier work this paper cites.
Practical Cryptography
Niels Ferguson and Bruce Schneier. 2003 · 2003
Earlier work this paper cites.
English gigaword
David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda. 2003 · 2003
Earlier work this paper cites.
The Enron corpus: A new dataset for email classification research
Bryan Klimt and Yiming Yang. 2004 · 2004
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Original
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2005
Earlier work this paper cites.
Language models are few-shot learners
Original
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Mt-adapted datasheets for datasets: Template and repository
Original
Marta R Costa-jussà, Roger Creus, Oriol Domingo, Albert Domínguez, Miquel Escobar, Cayetana López, Marina Garcia, and Margarita Geleta. 2020 · 2005
Earlier work this paper cites.
Social biases in NLP models as barriers for persons with disabilities
Original
Ben Hutchinson, Vinodkumar Prabhakaran, Emily Denton, Kellie Webster, Yu Zhong, and Stephen Denuyl. 2020 · 2005
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn. 2005 · 2005
Earlier work this paper cites.
Hybridity in mt: Experiments on the Europarl corpus
Declan Groves and Andy Way. 2006 · 2006
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Original
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020 · 2006
Earlier work this paper cites.
Large image datasets: A pyrrhic win for computer vision?
Original
Vinay Uday Prabhu and Abeba Birhane. 2020 · 2006
Earlier work this paper cites.
Source language markers in Europarl translations
Hans Van Halteren. 2008 · 2008
Earlier work this paper cites.
Sentiwordnet 3.0: An enhanced lexical resource for sentiment analysis and opinion mining
Stefano Baccianella, Andrea Esuli, and Fabrizio Sebastiani. 2010 · 2010
Earlier work this paper cites.
Language id in the wild: Unexpected challenges on the path to a thousand-language web text corpus
Original
Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna. 2020 · 2010
Earlier work this paper cites.