Fetching the paper…
Reading the bibliography…
While recent work has shown that scores from models trained by the ubiquitous masked language modeling (MLM) objective effectively discriminate probable from improbable sequences, it is still an open question if these MLMs specify a principled probability distribution over the space of possible sequences.
Monte carlo sampling methods using markov chains and their applications
W. K. Hastings · 1970
Earlier work this paper cites.
Sampling-based approaches to calculating marginal densities
A. E. Gelfand and A. F. Smith · 1990
Earlier work this paper cites.
Gibbs sampling
A. E. Gelfand · 2000
Earlier work this paper cites.
Dependency networks for inference, collaborative filtering, and data visualization
D. Heckerman, D. M. Chickering, C. Meek, R. Rounthwaite, and C. Kadie · 2000
Earlier work this paper cites.
Using conditional distributions for missing-data imputation
A. Gelman and T. E. Raghunathan · 2001
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning · 2003
Earlier work this paper cites.
Sparse graphical models for exploring gene expression data
A. Dobra, C. Hans, B. Jones, J. R. Nevins, G. Yao, and M. West · 2004
Earlier work this paper cites.
A tutorial on energy-based learning
Y. LeCun, S. Chopra, R. Hadsell, M. Ranzato, and F. Huang · 2006
Earlier work this paper cites.
Closed-form learning of markov networks from dependency networks
D. Lowd · 2012
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Structured prediction energy networks
D. Belanger and A. McCallum · 2016
Earlier work this paper cites.
Sequence-to-sequence learning as beam-search optimization
S. Wiseman and A. M. Rush · 2016
Earlier work this paper cites.
Energy-based generative adversarial network
J. Zhao, M. Mathieu, and Y. LeCun · 2016
Cited alongside, same era.
Non-autoregressive neural machine translation
J. Gu, J. Bradbury, C. Xiong, V. O. Li, and R. Socher · 2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Language modeling with neural trans-dimensional random fields
B. Wang and Z. Ou · 2017
Cited alongside, same era.
Adversarial feature matching for text generation
Y. Zhang, Z. Gan, K. Fan, Z. Chen, R. Henao, D. Shen, and L. Carin · 2017
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Later among the works it cites.
A generalized framework of sequence generation with application to undirected sequence models
E. Mansimov, A. Wang, S. Welleck, and K. Cho · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Later among the works it cites.
BERT has a mouth, and it must speak: BERT as a Markov random field language model
A. Wang and K. Cho · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Cited alongside, same era.
Learning neural trans-dimensional random field language models with noise-contrastive estimation
B. Wang and Z. Ou · 2018
Cited alongside, same era.
Implicit generation and modeling with energy based models
Y. Du and I. Mordatch · 2019
Cited alongside, same era.
Mask-predict: Parallel decoding of conditional masked language models
M. Ghazvininejad, O. Levy, Y. Liu, and L. Zettlemoyer · 2019
Cited alongside, same era.
An empirical investigation of global and local normalization for recurrent neural sequence models using a continuous relaxation to beam search
K. Goyal, C. Dyer, and T. Berg-Kirkpatrick · 2019
Cited alongside, same era.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Cited alongside, same era.
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V. Le · 2019
Later among the works it cites.
Bertscore: Evaluating text generation with BERT
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2019
Later among the works it cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Later among the works it cites.
Pre-training transformers as energy-based cloze models
K. Clark, M.-T. Luong, Q. Le, and C. D. Manning · 2020
Later among the works it cites.
Residual energy-based models for text generation
Y. Deng, A. Bakhtin, M. Ott, A. Szlam, and M. Ranzato · 2020
Later among the works it cites.
Engine: Energy-based inference networks for non-autoregressive machine translation
L. Tu, R. Y. Pang, S. Wiseman, and K. Gimpel · 2020
Later among the works it cites.
Oops I took a gradient: Scalable sampling for discrete distributions
W. Grathwohl, K. Swersky, M. Hashemi, D. Duvenaud, and C. J. Maddison · 2021
Closest in time.
A primer in Bertology: What we know about how BERT works
A. Rogers, O. Kovaleva, and A. Rumshisky · 2021
Closest in time.