Fetching the paper…
Reading the bibliography…
Masked language models (MLM) do not explicitly define a distribution over language, i.e., they are not language models per se.
Spatial interaction and the statistical analysis of lattice systems
Julian Besag. 1974 · 1974
Earlier work this paper cites.
Compatible conditional distributions
Barry C. Arnold and James S. Press. 1989 · 1989
Earlier work this paper cites.
Characterizing a joint probability distribution by conditionals
Andrew Gelman and Terence P. Speed. 1993 · 1993
Earlier work this paper cites.
Distributions most nearly compatible with given families of conditional distributions
Barry C. Arnold and Dattaprabhakar V. Gokhale. 1998 · 1998
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F. Chen and Joshua Goodman. 1998 · 1998
Earlier work this paper cites.
Dependency networks for inference, collaborative filtering, and data visualization
David Heckerman, Max Chickering, Chris Meek, Robert Rounthwaite, and Carl Kadie. 2000 · 2000
Earlier work this paper cites.
Exact and near compatibility of discrete conditional distributions
Barry C. Arnold, Enrique Castillo, and José María Sarabia. 2002 · 2002
Earlier work this paper cites.
Compatibility of finite discrete conditional distributions
Chwan-Chin Song, Lung-An Li, Chong-Hong Chen, Thomas J. Jiang, and Kun-Lin Kuo. 2010 · 2010
Earlier work this paper cites.
Compatibility of discrete conditional distributions with structural zeros
Yuchung J. Wang and Kun-Lin Kuo. 2010 · 2010
Earlier work this paper cites.
Closed-form learning of markov networks from dependency networks
Daniel Lowd. 2012 · 2012
Cited alongside, same era.
Generalized denoising auto-encoders as generative models
Yoshua Bengio, Li Yao, Guillaume Alain, and Pascal Vincent. 2013 · 2013
Cited alongside, same era.
Deep generative stochastic networks trainable by backprop
Yoshua Bengio, Éric Thibodeau-Laufer, Guillaume Alain, and Jason Yosinski. 2014 · 2014
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
BERT has a mouth, and it must speak: BERT as a Markov random field language model
Alex Wang and Kyunghyun Cho. 2019 · 2019
Later among the works it cites.
Masked language model scoring
Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Later among the works it cites.
A measure-theoretic characterization of tight language models
Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, and Ryan Cotterell. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Mask-Predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Exposing the implicit energy networks behind masked language models via Metropolis–Hastings
Kartik Goyal, Chris Dyer, and Taylor Berg-Kirkpatrick. 2022 · 2022
Later among the works it cites.
Probing BERT’s priors with serial reproduction chains
Takateru Yamakoshi, Thomas Griffiths, and Robert Hawkins. 2022 · 2022
Later among the works it cites.
On the inconsistencies of conditionals learned by masked language models
Tom Young and Yang You. 2023 · 2023
Closest in time.