Fetching the paper…
Reading the bibliography…
We introduce Electric, an energy-based cloze model for representation learning over text.
BERT has a mouth, and it must speak: BERT as a markov random field language model
Alex Wang and Kyunghyun Cho. 2019 · 1902
Earlier work this paper cites.
Implicit generation and generalization in energy-based models
Yilun Du and Igor Mordatch. 2019 · 1903
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
“Cloze procedure”: a new tool for measuring readability
Wilson L. Taylor. 1953 · 1953
Earlier work this paper cites.
The Helmholtz machine
Peter Dayan, Geoffrey E. Hinton, Radford M. Neal, and Richard S. Zemel. 1995 · 1995
Earlier work this paper cites.
Whole-sentence exponential language models: A vehicle for linguistic-statistical integration
Ronald Rosenfeld, Stanley F. Chen, and Xiaojin Zhu. 2001 · 2001
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2002
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E. Hinton. 2002 · 2002
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and William B. Dolan. 2007 · 2007
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu Jie Huang. 2007 · 2007
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Learning word embeddings efficiently with noise-contrastive estimation
Andriy Mnih and Koray Kavukcuoglu. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
On using very large target vocabulary for neural machine translation
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le. 2015 · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Trans-dimensional random fields for language modeling
Bin Wang, Zhijian Ou, and Zhiqiang Tan. 2015 · 2015
Cited alongside, same era.
A connection between generative adversarial networks, inverse reinforcement learning, and energy-based models
Chelsea Finn, Paul Christiano, Pieter Abbeel, and Sergey Levine. 2016 · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Józefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Cited alongside, same era.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy S. Liang. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Learning neural trans-dimensional random field language models with noise-contrastive estimation
Bin Wang and Zhijian Ou. 2018 · 2018
Later among the works it cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
Espnet: End-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junbo Zhao, Michael Mathieu, and Yann LeCun. 2016 · 2016
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel M. Cer, Mona T. Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
First Quora dataset release: Question pairs
Shankar Iyer, Nikhil Dandekar, and Kornél Csernai. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Learning with noise-contrastive estimation: Easing training by learning to scale
Matthieu Labeau and Alexandre Allauzen. 2018 · 2018
Cited alongside, same era.
Noise contrastive estimation and negative sampling for conditional models: Consistency and statistical efficiency
Zhuang Ma and Michael Collins. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Cloze-driven pretraining of self-attention networks
Alexei Baevski, Sergey Edunov, Yinhan Liu, Luke Zettlemoyer, and Michael Auli. 2019 · 2019
Later among the works it cites.
BAM! Born-again multi-task networks for natural language understanding
Kevin Clark, Minh-Thang Luong, Urvashi Khandelwal, Christopher D. Manning, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Effective sentence scoring method using bert for speech recognition
Joonbo Shin, Yoonhyung Lee, and Kyomin Jung. 2019 · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019 · 2019
Later among the works it cites.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
ELECTRA: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Closest in time.
Residual energy-based models for text generation
Yuntian Deng, Anton Bakhtin, Myle Ott, and Arthur Szlam. 2020 · 2020
Closest in time.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Closest in time.
Masked language model scoring
Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020 · 2020
Closest in time.