Fetching the paper…
Reading the bibliography…
We show that BERT (Devlin et al., 2018) is a Markov random field language model.
Cross-lingual Language Model Pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
Passage re-ranking with bert
Rodrigo Nogueira and Kyunghyun Cho. 2019 · 1901
Earlier work this paper cites.
Efficiency of pseudolikelihood estimation for simple gaussian fields
Julian Besag. 1977 · 1977
Earlier work this paper cites.
Replica monte carlo simulation of spin-glasses
Robert H Swendsen and Jian-Sheng Wang. 1986 · 1986
Earlier work this paper cites.
Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition
John S Bridle. 1990 · 1990
Earlier work this paper cites.
Probabilistic inference using markov chain monte carlo methods
Radford M Neal. 1993 · 1993
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Learning in markov random fields using tempered transitions
Ruslan R Salakhutdinov. 2009 · 2009
Earlier work this paper cites.
Parallel tempering is efficient for learning restricted boltzmann machines
KyungHyun Cho, Tapani Raiko, and Alexander Ilin. 2010 · 2010
Cited alongside, same era.
Tempered markov chain monte carlo for training of restricted boltzmann machines
Guillaume Desjardins, Aaron Courville, Yoshua Bengio, Pascal Vincent, and Olivier Delalleau. 2010 · 2010
Cited alongside, same era.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol. 2010 · 2010
Cited alongside, same era.
A connection between score matching and denoising autoencoders
Pascal Vincent. 2011 · 2011
Cited alongside, same era.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Later among the works it cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu. 2017 · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Later among the works it cites.
Multilingual constituency parsing with self-attention and pre-training
Nikita Kitaev and Dan Klein. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yacine Jernite, Alexander Rush, and David Sontag. 2015 · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Richard Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. 2016 · 2016
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Later among the works it cites.
Assessing bert’s syntactic abilities
Yoav Goldberg. 2019 · 2019
Closest in time.