Fetching the paper…
Reading the bibliography…
Energy-based models (EBMs), a.k.a.
Implicit generation and generalization in energy-based models
Yilun Du and Igor Mordatch · 1903
Earlier work this paper cites.
GLTR: statistical detection and visualization of generated text
Sebastian Gehrmann, Hendrik Strobelt, and Alexander M. Rush · 1906
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John Hopfield · 1982
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E. Schapire · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and K. Paliwal Kuldip · 1997
Earlier work this paper cites.
Robust real-time object detection
P. Viola and M. Jones · 2001
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E. Hinton · 2002
Earlier work this paper cites.
Energy-based models for sparse overcomplete representations
Y. W. Teh, M. Welling, S. Osindero, and Hinton G. E · 2003
Earlier work this paper cites.
Discriminative reranking for machine translation
Libin Shen, Anoop Sarkar, and Franz J. Och · 2004
Earlier work this paper cites.
Framewise phoneme classification with bidirectional lstm and other neural network architectures
A. Graves and J. Schmidhuber · 2005
Earlier work this paper cites.
Exponential family harmoniums with an application to information retrieval
Max Welling, Michal Rosen-Zvi, and Geoffrey E. Hinton · 2005
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu-Jie Huang · 2006
Cited alongside, same era.
A unified energy-based framework for unsupervised learning
Marc’Aurelio Ranzato, Y-Lan Boureau, Sumit Chopra, and Yann LeCun · 2007
Cited alongside, same era.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa · 2011
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Cited alongside, same era.
Modeling natural images using gated mrfs
M. Ranzato, V. Mnih, J. Susskind, and G.E. Hinton · 2013
Cited alongside, same era.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Later among the works it cites.
Importance of a search strategy in neural dialogue modelling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S Zemel, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Ilya Kulikov, Alexander H Miller, Kyunghyun Cho, and Jason Weston · 2018
Later among the works it cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2019
Closest in time.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Closest in time.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Closest in time.
URL https://github.com/openai/gpt-2-output-dataset/blob/master/README.md
Alec Radford and Jeff Wu, 2019 · 2019
Closest in time.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi · 2019
Closest in time.