Fetching the paper…
Reading the bibliography…
In this paper we propose a new generative model of text, Step-unrolled Denoising Autoencoder (SUNDAE), that does not rely on autoregressive models.
Neural networks and physical systems with emergent collective computational abilities
John J Hopfield · 1982
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E Hinton · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin · 2003
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning · 2003
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and F Huang · 2006
Earlier work this paper cites.
A unified energy-based framework for unsupervised learning
Marc’Aurelio Ranzato, Y-Lan Boureau, Sumit Chopra, and Yann LeCun · 2007
Earlier work this paper cites.
Generating text with recurrent neural networks
Ilya Sutskever, James Martens, and Geoffrey E Hinton · 2011
Earlier work this paper cites.
Pre-training transformers as energy-based cloze models
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, and Phillipp Koehn · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Generating sentences from a continuous space
Samuel R Bowman, Luke Vilnis, Oriol Vinyals, Andrew M Dai, Rafal Jozefowicz, and Samy Bengio · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
Danilo Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Density estimation using real nvp
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Maximum-likelihood augmented discrete generative adversarial networks
Tong Che, Yanran Li, Ruixiang Zhang, R Devon Hjelm, Wenjie Li, Yangqiu Song, and Yoshua Bengio · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher · 2017
Cited alongside, same era.
Long text generation via adversarial training with leaked information. arxiv e-prints, art
Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang, Yong Yu, and Jun Wang · 2017
Cited alongside, same era.
Adversarial ranking for language generation
Kevin Lin, Dianqi Li, Xiaodong He, Zhengyou Zhang, and Ming-Ting Sun · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Bert has a mouth, and it must speak: Bert as a markov random field language model
Alex Wang and Kyunghyun Cho · 2019
Later among the works it cites.
Non-autoregressive machine translation with auxiliary regularization
Yiren Wang, Fei Tian, Di He, Tao Qin, ChengXiang Zhai, and Tie-Yan Liu · 2019
Later among the works it cites.
Do sequence-to-sequence vaes learn global features of sentences?
Tom Bosc and Pascal Vincent · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Language gans falling short
Massimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle, Joelle Pineau, and Laurent Charlin · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Seqgan: Sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Fast decoding in sequence models using discrete latent variables
Lukasz Kaiser, Samy Bengio, Aurko Roy, Ashish Vaswani, Niki Parmar, Jakob Uszkoreit, and Noam Shazeer · 2018
Cited alongside, same era.
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Jason Lee, Elman Mansimov, and Kyunghyun Cho · 2018
Cited alongside, same era.
A call for clarity in reporting bleu scores
Matt Post · 2018
Cited alongside, same era.
Training language gans from scratch
Cyprien de Masson d’Autume, Mihaela Rosca, Jack Rae, and Shakir Mohamed · 2019
Cited alongside, same era.
Later among the works it cites.
Imputer: Sequence modelling via imputation and dynamic programming
William Chan, Chitwan Saharia, Geoffrey Hinton, Mohammad Norouzi, and Navdeep Jaitly · 2020
Later among the works it cites.
Residual energy-based models for text generation
Yuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam, and Marc’Aurelio Ranzato · 2020
Later among the works it cites.
Semi-autoregressive training improves mask-predict decoding
Marjan Ghazvininejad, Omer Levy, and Luke Zettlemoyer · 2020
Later among the works it cites.
Jointly masked sequence-to-sequence model for non-autoregressive neural machine translation
Junliang Guo, Linli Xu, and Enhong Chen · 2020
Later among the works it cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Later among the works it cites.
Non-autoregressive machine translation with disentangled context transformer
Jungo Kasai, James Cross, Marjan Ghazvininejad, and Jiatao Gu · 2020
Later among the works it cites.
Incorporating a local translation mechanism into non-autoregressive translation
Xiang Kong, Zhisong Zhang, and Eduard Hovy · 2020
Later among the works it cites.
Iterative refinement in the continuous space for non-autoregressive neural machine translation
Jason Lee, Raphael Shu, and Kyunghyun Cho · 2020
Later among the works it cites.
Non-autoregressive machine translation with latent alignments
Chitwan Saharia, William Chan, Saurabh Saxena, and Mohammad Norouzi · 2020
Later among the works it cites.
Latent-variable non-autoregressive neural machine translation with deterministic inference using a delta posterior
Raphael Shu, Jason Lee, Hideki Nakayama, and Kyunghyun Cho · 2020
Later among the works it cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2020
Later among the works it cites.
Structured denoising diffusion models in discrete state-spaces
Jacob Austin, Daniel Johnson, Jonathan Ho, Danny Tarlow, and Rianne van den Berg · 2021
Closest in time.
Argmax flows and multinomial diffusion: Towards non-autoregressive language models
Emiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré, and Max Welling · 2021
Closest in time.
Text 8 dataset
Matt Mahoney · 2021
Closest in time.
Symbolic music generation with diffusion models
Gautam Mittal, Jesse Engel, Curtis Hawthorne, and Ian Simon · 2021
Closest in time.
Improved denoising diffusion probabilistic models
Alex Nichol and Prafulla Dhariwal · 2021
Closest in time.
Maskgit: Masked generative image transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman · 2022
Closest in time.