Fetching the paper…
Reading the bibliography…
The predominant approach for language modeling is to process sequences from left to right, but this eliminates a source of information: the order by which the sequence was generated.
Non-monotonic sequential text generation
Sean Welleck, Kianté Brantley, Hal Daumé III, and Kyunghyun Cho · 1902
Earlier work this paper cites.
Algorithms for the Assignment and Transportation Problems
James R. Munkres · 1957
Earlier work this paper cites.
Individual Choice Behavior: A Theoretical analysis
R. Duncan Luce · 1959
Earlier work this paper cites.
A relationship between arbitrary positive matrices and doubly stochastic matrices
Richard Sinkhorn · 1964
Earlier work this paper cites.
The analysis of permutations
Robin L Plackett · 1975
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Insertion-deletion transformer
Laura Ruis, Mitchell Stern, Julia Proskurnia, and William Chan · 2001
Earlier work this paper cites.
A syntax-based statistical translation model
Kenji Yamada and Kevin Knight · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Syntax-based language models for statistical machine translation
Eugene Charniak, Kevin Knight, and Kenji Yamada · 2003
Earlier work this paper cites.
English gigaword, 2003
David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul · 2006
Earlier work this paper cites.
The bethe permanent of a non-negative matrix
P. O. Vontobel · 2010
Earlier work this paper cites.
Generating text with recurrent neural networks
Ilya Sutskever, James Martens, and Geoffrey E Hinton · 2011
Earlier work this paper cites.
Statistical language models based on neural networks
Tomáš Mikolov et al · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Made: Masked autoencoder for distribution estimation
Mathieu Germain, Karol Gregor, Iain Murray, and Hugo Larochelle · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Microsoft coco: Common objects in context, 2015
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Cited alongside, same era.
Learning to generate pseudo-code from source code using statistical machine translation
Yusuke Oda, Hiroyuki Fudaba, Graham Neubig, Hideaki Hata, Sakriani Sakti, Tomoki Toda, and Satoshi Nakamura · 2015
Cited alongside, same era.
Pixelsnail: An improved autoregressive generative model
Xi Chen, Nikhil Mishra, Mostafa Rohaninejad, and Pieter Abbeel · 2018
Later among the works it cites.
Top-down tree structured decoding with syntactic connections for neural machine translation and parsing
Jetic Gū, Hassan S. Shavarani, and Anoop Sarkar · 2018
Later among the works it cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor O. K. Li, and Richard Socher · 2018
Later among the works it cites.
Reparameterizing the birkhoff polytope for variational permutation inference
Scott Linderman, Gonzalo Mena, Hal Cooper, Liam Paninski, and John Cunningham · 2018
Later among the works it cites.
Middle-out decoding
Shikib Mehri and Leonid Sigal · 2018
Later among the works it cites.
Learning latent permutations with gumbel-sinkhorn networks
Gonzalo Mena, David Belanger, Scott Linderman, and Jasper Snoek · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A neural attention model for abstractive sentence summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
Pointer networks
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A Smith · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush · 2016
Cited alongside, same era.
Later among the works it cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Later among the works it cites.
A tree-based decoder for neural machine translation
Xinyi Wang, Hieu Pham, Pengcheng Yin, and Graham Neubig · 2018
Later among the works it cites.
Beyond error propagation in neural machine translation: Characteristics of language also matter
Lijun Wu, Xu Tan, Di He, Fei Tian, Tao Qin, Jianhuang Lai, and Tie-Yan Liu · 2018
Later among the works it cites.
A tight analysis of bethe approximation for permanent
Nima Anari and Alireza Rezaei · 2019
Later among the works it cites.
Kermit: Generative insertion-based modeling for sequences, 2019
William Chan, Nikita Kitaev, Kelvin Guu, Mitchell Stern, and Jakob Uszkoreit · 2019
Later among the works it cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Sequence modeling with unconstrained generation order
Dmitrii Emelianenko, Elena Voita, and Pavel Serdyukov · 2019
Later among the works it cites.
Stochastic optimization of sorting networks via continuous relaxations
Aditya Grover, E. Wang, Aaron Zweig, and S. Ermon · 2019
Later among the works it cites.
Levenshtein transformer
Jiatao Gu, Changhan Wang, and Junbo Zhao · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Neural machine translation: A review
Felix Stahlberg · 2019
Later among the works it cites.
Insertion transformer: Flexible sequence generation via insertion operations
Mitchell Stern, William Chan, Jamie Kiros, and Jakob Uszkoreit · 2019
Later among the works it cites.
Non-monotonic sequential text generation
Sean Welleck, Kianté Brantley, Hal Daumé III, and Kyunghyun Cho · 2019
Later among the works it cites.
Synchronous bidirectional neural machine translation
Long Zhou, Jiajun Zhang, and Chengqing Zong · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Sinkhorn permutation variational marginal inference
Gonzalo Mena, Erdem Varol, Amin Nejatbakhsh, Eviatar Yemini, and Liam Paninski · 2020
Later among the works it cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2074
Closest in time.