Fetching the paper…
Reading the bibliography…
Exposure bias has been regarded as a central problem for auto-regressive language models (LM).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Training language gans from scratch
Cyprien de Masson d’Autume, Mihaela Rosca, Jack W. Rae, and Shakir Mohamed. 2019 · 1905
Earlier work this paper cites.
Bridging the gap between training and inference for neural machine translation
Wen Zhang, Yang Feng, Fandong Meng, Di You, and Qun Liu. 2019 · 1906
Earlier work this paper cites.
Rethinking exposure bias in language modeling
Yifan Xu, Kening Zhang, Haoyu Dong, Yuezhou Sun, Wenlong Zhao, and Zhuowen Tu. 2019 · 1910
Earlier work this paper cites.
Negated LAMA: birds cannot fly
Nora Kassner and Hinrich Schütze. 2019 · 1911
Earlier work this paper cites.
How decoding strategies affect the verifiability of generated text
Luca Massarelli, Fabio Petroni, Aleksandra Piktus, Myle Ott, Tim Rocktäschel, Vassilis Plachouras, Fabrizio Silvestri, and Sebastian Riedel. 2019 · 1911
Earlier work this paper cites.
Trainable greedy decoding for neural machine translation
Jiatao Gu, Kyunghyun Cho, and Victor O.K. Li. 2017 · 1978
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning , 1st edition
Richard S. Sutton and Andrew G. Barto. 1998 · 1998
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
LSTM neural networks for language modeling
Martin Sundermeyer, Ralf Schlüter, and Hermann Ney. 2012 · 2012
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie. 2014 · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Earlier work this paper cites.
How (not) to train your generative model: Scheduled sampling, likelihood, adversary?
Ferenc Huszár. 2015 · 2015
Earlier work this paper cites.
Professor forcing: A new algorithm for training recurrent networks
Alex Lamb, Anirudh Goyal, Ying Zhang, Saizheng Zhang, Aaron C. Courville, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
Self-critical sequence training for image captioning
Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. 2016 · 2016
Cited alongside, same era.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M. Rush. 2016 · 2016
Cited alongside, same era.
Seqgan: Sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu. 2016 · 2016
On accurate evaluation of gans for language generation
Stanislau Semeniuta, Aliaksei Severyn, and Sylvain Gelly. 2018 · 2018
Later among the works it cites.
Towards diverse text generation with inverse reinforcement learning
Zhan Shi, Xinchi Chen, Xipeng Qiu, and Xuanjing Huang. 2018 · 2018
Later among the works it cites.
Evaluating text gans as language models
Guy Tevet, Gavriel Habib, Vered Shwartz, and Jonathan Berant. 2018 · 2018
Later among the works it cites.
Generating informative and diverse conversational responses via adversarial information maximization
Yizhe Zhang, Michel Galley, Jianfeng Gao, Zhe Gan, Xiujun Li, Chris Brockett, and Bill Dolan. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Long text generation via adversarial training with leaked information
Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang, Yong Yu, and Jun Wang. 2017 · 2017
Cited alongside, same era.
Adversarial learning for neural dialogue generation
Jiwei Li, Will Monroe, Tianlin Shi, Alan Ritter, and Dan Jurafsky. 2017 · 2017
Cited alongside, same era.
Adversarial ranking for language generation
Kevin Lin, Dianqi Li, Xiaodong He, Zhengyou Zhang, and Ming-ting Sun. 2017 · 2017
Cited alongside, same era.
Adversarial generation of natural language
Sai Rajeswar, Sandeep Subramanian, Francis Dutil, Christopher Joseph Pal, and Aaron C. Courville. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Adversarial neural machine translation
Lijun Wu, Yingce Xia, Li Zhao, Fei Tian, Tao Qin, Jianhuang Lai, and Tie-Yan Liu. 2017 · 2017
Cited alongside, same era.
Recent trends in deep learning based natural language processing
Tom Young, Devamanyu Hazarika, Soujanya Poria, and Erik Cambria. 2017 · 2017
Cited alongside, same era.
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Later among the works it cites.
Ctrl: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019 · 2019
Closest in time.
RelGAN: Relational generative adversarial networks for text generation
Weili Nie, Nina Narodytska, and Ankit Patel. 2019 · 2019
Closest in time.
Generalization in generation: A closer look at exposure bias
Florian Schmidt. 2019 · 2019
Closest in time.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Closest in time.
A systematic characterization of sampling algorithms for open-ended language generation
Moin Nadeem, Tianxing He, Kyunghyun Cho, and James Glass. 2020 · 2020
Closest in time.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Closest in time.
On exposure bias, hallucination and domain shift in neural machine translation
Chaojun Wang and Rico Sennrich. 2020 · 2020
Closest in time.
Neural text generation with unlikelihood training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020 · 2020
Closest in time.
Trading off diversity and quality in natural language generation
Hugh Zhang, Daniel Duckworth, Daphne Ippolito, and Arvind Neelakantan. 2020 · 2020
Closest in time.
Guiding teacher forcing with seer forcing for neural machine translation
Yang Feng, Shuhao Gu, Dengji Guo, Zhengxin Yang, and Chenze Shao. 2021 · 2021
Closest in time.
Analyzing the forgetting problem in pretrain-finetuning of open-domain dialogue response models
Tianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James Glass, and Fuchun Peng. 2021 · 2021
Closest in time.
Negative training for neural dialogue response generation
Tianxing He and James Glass. 2020 · 2058
Closest in time.