Fetching the paper…
Reading the bibliography…
This datasheet describes the Pile, a 825 GiB dataset of human-authored text compiled by EleutherAI for use in large-scale language modeling.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap · 1911
Earlier work this paper cites.
The Enron corpus: A new dataset for email classification research
Bryan Klimt and Yiming Yang · 2004
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn · 2005
Earlier work this paper cites.
Communication networks from the enron email corpus “it’s always about the people. enron is no different”
Jana Diesner, Terrill L Frantz, and Kathleen M Carley · 2005
Earlier work this paper cites.
Hybridity in mt: Experiments on the Europarl corpus
Declan Groves and Andy Way · 2006
Earlier work this paper cites.
Timeline: Enron, Jan 2006
Staff and agencies · 2006
Earlier work this paper cites.
Source language markers in Europarl translations
Hans Van Halteren · 2008
Earlier work this paper cites.
The structure of information pathways in a social communication network
Gueorgi Kossinets, Jon Kleinberg, and Duncan Watts · 2008
Earlier work this paper cites.
Community evolution in dynamic multi-mode networks
Lei Tang, Huan Liu, Jianping Zhang, and Zohreh Nazeri · 2008
Earlier work this paper cites.
The reach and richness of wikipedia: Is wikinomics only for rich countries?
Morten Rask · 2008
Earlier work this paper cites.
Normalized (pointwise) mutual information in collocation extraction
Gerlof Bouma · 2009
Earlier work this paper cites.
Statistical machine translation
Philipp Koehn · 2009
Earlier work this paper cites.
Adaptive regularization of weight vectors
Koby Crammer, Alex Kulesza, and Mark Dredze · 2009
Earlier work this paper cites.
Kenlm: Faster and smaller language model queries
Kenneth Heafield · 2011
Earlier work this paper cites.
An extensive experimental comparison of methods for multi-label learning
Gjorgji Madjarov, Dragi Kocev, Dejan Gjorgjevikj, and Sašo Džeroski · 2011
Earlier work this paper cites.
Gender bias in wikipedia and britannica
Joseph Reagle and Lauren Rhue · 2011
Earlier work this paper cites.
Cultural bias in wikipedia content on famous persons
Ewa S Callahan and Susan C Herring · 2011
Earlier work this paper cites.
Babelnet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network
Roberto Navigli and Simone Paolo Ponzetto · 2012
Earlier work this paper cites.
Art education and disability studies
John Derby · 2012
Earlier work this paper cites.
More effective boilerplate removal – the GoldMiner algorithm
István Endrédy and Attila Novák · 2013
Earlier work this paper cites.
The Ubuntu chat corpus for multiparticipant chat analysis
David C Uthus and David W Aha · 2013
Earlier work this paper cites.
Bi: Notes for a bisexual revolution
Shiri Eisner · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation, 2014
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Russ R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
The enron data set - where did it come from?
Joe Bartling · 2015
Earlier work this paper cites.
It’s a man’s wikipedia? assessing gender inequality in an online encyclopedia
Claudia Wagner, David Garcia, Mohsen Jadidi, and Markus Strohmaier · 2015
Cited alongside, same era.
Mind the skills gap: the role of internet know-how and gender in differentiated contributions to wikipedia
Eszter Hargittai and Aaron Shaw · 2015
Cited alongside, same era.
First women, second sex: Gender bias in wikipedia
Eduardo Graells-Garrido, Mounia Lalmas, and Filippo Menczer · 2015
Cited alongside, same era.
Finding alternative translations in a large corpus of movie subtitles
J. Tiedemann · 2016
Cited alongside, same era.
The ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems, 2016
Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau · 2016
Cited alongside, same era.
Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions
Yu-Siang Huang and Yi-Hsuan Yang · 2020
Later among the works it cites.
How much self-attention do we need? trading attention for feed-forward layers
K. Irie, A. Gerstenberger, R. Schlüter, and H. Ney · 2020
Later among the works it cites.
Distill, adapt, distill: Training small, in-domain models for neural machine translation, 2020
Mitchell A. Gordon and Kevin Duh · 2020
Later among the works it cites.
Tilde at wmt 2020: News task systems
Rihards Krišlauks and Mārcis Pinnis · 2020
Later among the works it cites.
Leap-Of-Thought: Teaching pre-trained models to systematically reason over implicit knowledge
Alon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg, and Jonathan Berant · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Cited alongside, same era.
Layer normalization, 2016
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Cited alongside, same era.
Wikipedia and history: a worthwhile partnership in the digital era?
Murray G Phillips · 2016
Cited alongside, same era.
Towards a map of the syntactic similarity of languages
Alina Maria Ciobanu, Liviu P Dinu, and Andrea Sgarro · 2017
Cited alongside, same era.
Information Extraction from TV Series Scripts for Uptake Prediction
Junshu Wang · 2017
Cited alongside, same era.
Nation image and its dynamic changes in wikipedia
Youngwhan Lee and Heuiju Chun · 2017
Cited alongside, same era.
Paraphrase detection on noisy subtitles in six languages, 2018
Eetu Sjöblom, Mathias Creutz, and Mikko Aulamo · 2018
Cited alongside, same era.
Joint translation and unit conversion for end-to-end localization, 2020
Georgiana Dinu, Prashant Mathur, Marcello Federico, Stanislas Lauly, and Yaser Al-Onaizan · 2020
Later among the works it cites.
Performance vs. competence in human–machine comparisons
Chaz Firestone · 2020
Later among the works it cites.
URL https://www.ncbi.nlm.nih.gov/pmc/about/faq/
2020 · 2020
Later among the works it cites.
Editing for equity: Understanding instructor motivations for integrating cross-disciplinary wikipedia assignments
Jiawei Xing and Matthew Vetter · 2020
Later among the works it cites.
Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Later among the works it cites.
Gpt-j-6b: A 6 billion parameter autoregressive language model, 2021
Ben Wang and Aran Komatsuzaki · 2021
Later among the works it cites.
Jurassic-1: Technical details and evaluation
Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham · 2021
Later among the works it cites.
WuDao: pretrain the world
Jie Tang · 2021
Later among the works it cites.
GPT-NeoX: Large scale autoregressive language modeling in pytorch, 2021
Alex Andonian, Quentin Anthony, Stella Biderman, Sid Black, Preetham Gali, Leo Gao, Eric Hallahan, Josh Levy-Kramer, Connor Leahy, Lucas Nestler, Kip Parker, Michael Pieler, Shivanshu Purohit, Tri Songz, Phil Wang, and Samuel Weinbach · 2021
Later among the works it cites.
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning · 2021
Later among the works it cites.
Maxime Peyrard, Sarvjeet Singh Ghotra, Martin Josifoski, Vidhan Agarwal, Barun Patra, Dean Carignan, Emre Kiciman, and Robert West · 2021
Later among the works it cites.
Cut the carp: Fishing for zero-shot story evaluation
Shahbuland Matiana, JR Smith, Ryan Teehan, Louis Castricato, Stella Biderman, Leo Gao, and Spencer Frazier · 2021
Later among the works it cites.
Neural program generation modulo static analysis
Rohan Mukherjee, Yeming Wen, Dipak Chaudhari, Thomas Reps, Swarat Chaudhuri, and Chris Jermaine · 2021
Later among the works it cites.
Intersectional bias in causal language models
Liam Magee, Lida Ghahremanlou, Karen Soldatic, and Shanthi Robertson · 2021
Later among the works it cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2021
Later among the works it cites.
Analyzing the implicit position encoding ability of transformer decoder
Ziyang Luo, Yadong Xi, Jing Ma, Xiaoxi Mao, and Changjie Fan · 2021
Later among the works it cites.
Using DeepSpeed and Megatron to train Megatron-Turing NLG 530B, the world’s largest and most powerful generative language model, Oct 2021
Paresh Kharya and Ali Alvi · 2021
Later among the works it cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al · 2021
Later among the works it cites.
Phrase-bert: Improved phrase embeddings from bert with an application to corpus exploration
Shufan Wang, Laure Thompson, and Mohit Iyyer · 2021
Later among the works it cites.
Jack Bandy and Nicholas Vincent · 2021
Later among the works it cites.
Using wikipedia to explore issues of systemic bias and symbolic annihilation in information sources
Caroline Ball · 2021
Later among the works it cites.