Fetching the paper…
Reading the bibliography…
Many NLP applications, such as biomedical data and technical support, have 10-100 million tokens of in-domain data and limited computational resources for learning from it.
Single headed attention rnn: Stop thinking with your head
Stephen Merity. 2019 · 1911
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
North american news text corpus LDC95T21
David Graff. 1995 · 1995
Earlier work this paper cites.
Train no evil: Selective masking for task-guided pre-training
Yuxian Gu, Zhengyan Zhang, Xiaozhi Wang, Zhiyuan Liu, and Maosong Sun. 2020 · 2004
Earlier work this paper cites.
Scalable modified Kneser-Ney language model estimation
Kenneth Heafield, Ivan Pouzyrevsky, Jonathan H. Clark, and Philipp Koehn. 2013 · 2013
Earlier work this paper cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Learning to create and reuse words in open-vocabulary neural language modeling
Kazuya Kawakami, Chris Dyer, and Phil Blunsom. 2017 · 2017
Earlier work this paper cites.
An exploration of word embedding initialization in deep-learning tasks
Tom Kocmi and Ondřej Bojar. 2017 · 2017
Earlier work this paper cites.
A bag of useful tricks for practical neural machine translation: Embedding layer initialization and large batch size
Masato Neishi, Jin Sakuma, Satoshi Tohda, Shonosuke Ishiwatari, Naoki Yoshinaga, and Masashi Toyoda. 2017 · 2017
Earlier work this paper cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf. 2017 · 2017
Earlier work this paper cites.
Unsupervised neural machine translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018 · 2018
Cited alongside, same era.
Are all languages equally hard to language-model?
Ryan Cotterell, Sabrina J. Mielke, Jason Eisner, and Brian Roark. 2018 · 2018
Cited alongside, same era.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom. 2018 · 2018
Cited alongside, same era.
Beyond weight tying: Learning joint input-output embeddings for neural machine translation
Nikolaos Pappas, Lesly Miculicich, and James Henderson. 2018 · 2018
Cited alongside, same era.
When and why are pre-trained word embeddings useful for neural machine translation?
Ye Qi, Devendra Sachan, Matthieu Felix, Sarguna Padmanabhan, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Show some love to your n-grams: A bit of progress and stronger n-gram language modeling baselines
Ehsan Shareghi, Daniela Gerz, Ivan Vulić, and Anna Korhonen. 2019 · 2019
Later among the works it cites.
Contextual embeddings: When are they worth it?
Simran Arora, Avner May, Jian Zhang, and Christopher Ré. 2020 · 2020
Closest in time.
Stolen probability: A structural weakness of neural language models
David Demeter, Gregory Kimmel, and Doug Downey. 2020 · 2020
Closest in time.
Exploiting syntactic structure for better language modeling: A syntactic distance approach
Wenyu Du, Zhouhan Lin, Yikang Shen, Timothy J. O’Donnell, Yoshua Bengio, and Yue Zhang. 2020 · 2020
Closest in time.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Representation degeneration problem in training natural language generation models
Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2019 · 2019
Cited alongside, same era.
A large-scale corpus for conversation disentanglement
Jonathan K. Kummerfeld, Sai R. Gouravajhala, Joseph J. Peper, Vignesh Athreya, Chulaka Gunasekara, Jatin Ganhotra, Siva Sankalp Patel, Lazaros Polymenakos, and Walter S. Lasecki. 2019 · 2019
Cited alongside, same era.
ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing
Mark Neumann, Daniel King, Iz Beltagy, and Waleed Ammar. 2019 · 2019
Cited alongside, same era.
To tune or not to tune? adapting pretrained representations to diverse tasks
Matthew E. Peters, Sebastian Ruder, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Regularizing and Optimizing LSTM Language Models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2017a
Cited in the paper.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017b
Cited in the paper.
Yinqiao Li, Chi Hu, Yuhao Zhang, Nuo Xu, Yufan Jiang, Tong Xiao, Jingbo Zhu, Tongran Liu, and Changliang Li. 2020 · 2020
Closest in time.
Grounded compositional outputs for adaptive language modeling
Nikolaos Pappas and Noah A. Mulcaire, Phoebe Smith. 2020 · 2020
Closest in time.
Stanza: A python natural language processing toolkit for many human languages
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020 · 2020
Closest in time.
BERTRAM: Improved word embeddings have big impact on contextualized model performance
Timo Schick and Hinrich Schütze. 2020 · 2020
Closest in time.
Cord-19: The covid-19 open research dataset
Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Darrin Eide, Kathryn Funk, Rodney Michael Kinney, Ziyang Liu, William. Merrill, Paul Mooney, Dewey A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D. Wade, Kuansan Wang, Christopher Wilhelm, Boya Xie, Douglas M. Raymond, Daniel S. Weld, Oren Etzioni, and Sebastian Kohlmeier. 2020 · 2020
Closest in time.