Fetching the paper…
Reading the bibliography…
Language model pre-training, such as BERT, has achieved remarkable results in many NLP tasks.
Assessing BERT’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Multi-task deep neural networks for natural language understanding
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019b · 1901
Earlier work this paper cites.
Chiyuan Zhang, Samy Bengio, and Yoram Singer. 2019 · 1902
Earlier work this paper cites.
Cloze-driven pretraining of self-attention networks
Alexei Baevski, Sergey Edunov, Yinhan Liu, Luke Zettlemoyer, and Michael Auli. 2019 · 1903
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew Peters, and Noah A. Smith. 2019a · 1903
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 1905
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, and Danilo Giampiccolo. 2006 · 2006
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006 · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007 · 2007
Earlier work this paper cites.
The fifth PASCAL recognizing textual entailment challenge
Luisa Bentivogli, Ido Dagan, Hoa Trang Dang, Danilo Giampiccolo, and Bernardo Magnini. 2009 · 2009
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Qualitatively characterizing neural network optimization problems
Ian J. Goodfellow and Oriol Vinyals. 2015 · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
An empirical analysis of deep network loss surfaces
Daniel Jiwoong Im, Michael Tao, and Kristin Branson. 2016 · 2016
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. 2018 · 2018
Later among the works it cites.
LSTMs can learn syntax-sensitive dependencies well, but modeling structure makes them better
Adhiguna Kuncoro, Chris Dyer, John Hale, Dani Yogatama, Stephen Clark, and Phil Blunsom. 2018 · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2016 · 2016
Cited alongside, same era.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
Entropy-SGD: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina. 2017 · 2017
Cited alongside, same era.
Learned in translation: Contextualized word vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018a
Cited in the paper.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018b
Cited in the paper.
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
To tune or not to tune? Adapting pretrained representations to diverse tasks
Matthew Peters, Sebastian Ruder, and Noah A. Smith. 2019 · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Closest in time.