Fetching the paper…
Reading the bibliography…
This work focuses on relating two mysteries in neural-based text generation: exposure bias, and text degeneration.
Liii. on lines and planes of closest fit to systems of points in space
Karl Pearson. 1901 · 1901
Earlier work this paper cites.
Quantifying exposure bias for neural language generation
Tianxing He, Jingzhao Zhang, Zhiming Zhou, and James Glass. 2019 · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Don’t say that! making inconsistent dialogue unlikely with unlikelihood training
Margaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck, Y-Lan Boureau, Kyunghyun Cho, and Jason Weston. 2019 · 1911
Earlier work this paper cites.
Dialogpt: Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2019b · 1911
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau. 1989 · 1989
Earlier work this paper cites.
Learning to play the game of chess
Sebastian Thrun. 1995 · 1995
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Stefan Schaal. 1999 · 1999
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Off-road obstacle avoidance through end-to-end learning
Urs Muller, Jan Ben, Eric Cosatto, Beat Flepp, and Yann L Cun. 2006 · 2006
Earlier work this paper cites.
Maximum margin planning
Nathan D Ratliff, J Andrew Bagnell, and Martin A Zinkevich. 2006 · 2006
Earlier work this paper cites.
Search-based structured prediction
Hal Daumé, John Langford, and Daniel Marcu. 2009 · 2009
Cited alongside, same era.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell. 2010 · 2010
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell. 2011 · 2011
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Cited alongside, same era.
How (not) to train your generative model: Scheduled sampling, likelihood, adversary?
Ferenc Huszár. 2015 · 2015
Cited alongside, same era.
Professor forcing: A new algorithm for training recurrent networks
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford. 2018 · 2018
Later among the works it cites.
Seq2seq-vis: A visual debugging tool for sequence-to-sequence models
Hendrik Strobelt, Sebastian Gehrmann, Michael Behrisch, Adam Perer, Hanspeter Pfister, and Alexander M Rush. 2018 · 2018
Later among the works it cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Later among the works it cites.
Generalization in generation: A closer look at exposure bias
Florian Schmidt. 2019 · 2019
Later among the works it cites.
Do massively pretrained language models make better storytellers?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex M Lamb, Anirudh Goyal Alias Parth Goyal, Ying Zhang, Saizheng Zhang, Aaron C Courville, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M. Rush. 2016 · 2016
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Seqgan: Sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu. 2017 · 2017
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever
Cited in the paper.
Abigail See, Aneesh Pappu, Rohun Saxena, Akhila Yerukola, and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Later among the works it cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Later among the works it cites.
On exposure bias, hallucination and domain shift in neural machine translation
Chaojun Wang and Rico Sennrich. 2020 · 2020
Later among the works it cites.
Neural text generation with unlikelihood training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020 · 2020
Later among the works it cites.