Fetching the paper…
Reading the bibliography…
The dominant language modeling paradigm handles text as a sequence of discrete tokens.
KERMIT: Generative insertion-based modeling for sequences
William Chan, Nikita Kitaev, Kelvin Guu, Mitchell Stern, and Jakob Uszkoreit. 2019 · 1906
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J Williams and David Zipser. 1989 · 1989
Earlier work this paper cites.
Structure and performance of a dependency language model
Ciprian Chelba, David Engle, Frederick Jelinek, Victor Jimenez, Sanjeev Khudanpur, Lidia Mangu, Harry Printz, Eric Ristad, Ronald Rosenfeld, Andreas Stolcke, and Dekai Wu. 1997 · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Nltk: The natural language toolkit
Edward Loper and Steven Bird. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
A new string-to-dependency machine translation algorithm with a target dependency language model
Libin Shen, Jinxi Xu, and Ralph Weischedel. 2008 · 2008
Earlier work this paper cites.
Dependency recurrent neural language models for sentence completion
Piotr Mirowski and Andreas Vlachos. 2015 · 2015
Earlier work this paper cites.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Neural syntactic generative models with exact marginalization
Jan Buys and Phil Blunsom. 2018 · 2018
Cited alongside, same era.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Jason Lee, Elman Mansimov, and Kyunghyun Cho. 2018 · 2018
Cited alongside, same era.
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Neural language modeling by jointly learning syntax and lexicon
Yikang Shen, Zhouhan Lin, Chin wei Huang, and Aaron Courville. 2018 · 2018
Mask-predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. 2019 · 2019
Later among the works it cites.
Attending to future tokens for bidirectional sequence generation
Carolin Lawrence, Bhushan Kotnis, and Mathias Niepert. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Ordered neurons: Integrating tree structures into recurrent neural networks
Yikang Shen, Shawn Tan, Alessandro Sordoni, and Aaron Courville. 2019 · 2019
Later among the works it cites.
Insertion transformer: Flexible sequence generation via insertion operations
Mitchell Stern, William Chan, Jamie Kiros, and Jakob Uszkoreit. 2019 · 2019
Later among the works it cites.
Non-monotonic sequential text generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Cited alongside, same era.
Syntactically supervised transformers for faster neural machine translation
Nader Akoury, Kalpesh Krishna, and Mohit Iyyer. 2019 · 2019
Cited alongside, same era.
Sequence modeling with unconstrained generation order
Dmitrii Emelianenko, Elena Voita, and Pavel Serdyukov. 2019 · 2019
Cited alongside, same era.
Insertion-based decoding with automatically inferred generation order
Jiatao Gu, Qi Liu, and Kyunghyun Cho. 2019a
Cited in the paper.
Levenshtein transformer
Jiatao Gu, Changhan Wang, and Junbo Zhao. 2019b
Cited in the paper.
Sean Welleck, Kianté Brantley, Hal Daumé III, and Kyunghyun Cho. 2019 · 2019
Later among the works it cites.
Language gans falling short
Massimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle, Joelle Pineau, and Laurent Charlin. 2020 · 2020
Closest in time.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Closest in time.