Fetching the paper…
Reading the bibliography…
We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers.
Semi-supervised sequence modeling with cross-view training
Kevin Clark, Minh-Thang Luong, Christopher D Manning, and Quoc Le. 2018 · 1925
Earlier work this paper cites.
“Cloze procedure”: A new tool for measuring readability
Wilson L Taylor. 1953 · 1953
Earlier work this paper cites.
Class-based n-gram models of natural language
Peter F Brown, Peter V Desouza, Robert L Mercer, Vincent J Della Pietra, and Jenifer C Lai. 1992 · 1992
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik F Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
A framework for learning predictive structures from multiple tasks and unlabeled data
Rie Kubota Ando and Tong Zhang. 2005 · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Domain adaptation with structural correspondence learning
John Blitzer, Ryan McDonald, and Fernando Pereira. 2006 · 2006
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Ronan Collobert and Jason Weston. 2008 · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. 2008 · 2008
Earlier work this paper cites.
The fifth PASCAL recognizing textual entailment challenge
Luisa Bentivogli, Bernardo Magnini, Ido Dagan, Hoa Trang Dang, and Danilo Giampiccolo. 2009 · 2009
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. 2009 · 2009
Earlier work this paper cites.
A scalable hierarchical distributed language model
Andriy Mnih and Geoffrey E Hinton. 2009 · 2009
Earlier work this paper cites.
Word representations: A simple and general method for semi-supervised learning
Joseph Turian, Lev Ratinov, and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
The winograd schema challenge
Hector J Levesque, Ernest Davis, and Leora Morgenstern. 2011 · 2011
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Distributed representations of sentences and documents
Quoc Le and Tomas Mikolov. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Cited alongside, same era.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. 2014 · 2014
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le. 2015 · 2015
Cited alongside, same era.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Bidirectional attention flow for machine comprehension
Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Contextual string embeddings for sequence labeling
Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018 · 2018
Closest in time.
Character-level language modeling with deeper self-attention
Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones. 2018 · 2018
Closest in time.
Quora question pairs
Z. Chen, H. Zhang, X. Zhang, and L. Zhao. 2018 · 2018
Closest in time.
Simple and effective multi-paragraph reading comprehension
Christopher Clark and Matt Gardner. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Cited alongside, same era.
Learning distributed representations of sentences from unlabelled data
Felix Hill, Kyunghyun Cho, and Anna Korhonen. 2016 · 2016
Cited alongside, same era.
context2vec: Learning generic context embedding with bidirectional LSTM
Oren Melamud, Jacob Goldberger, and Ido Dagan. 2016 · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
Ankur P Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Closest in time.
Maskgan: Better text generation via filling in the_
William Fedus, Ian Goodfellow, and Andrew M Dai. 2018 · 2018
Closest in time.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Closest in time.
Reinforced mnemonic reader for machine reading comprehension
Minghao Hu, Yuxing Peng, Zhen Huang, Xipeng Qiu, Furu Wei, and Ming Zhou. 2018 · 2018
Closest in time.
An efficient framework for learning sentence representations
Lajanugen Logeswaran and Honglak Lee. 2018 · 2018
Closest in time.
Dissecting contextual word embeddings: Architecture and representation
Matthew Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih. 2018b · 2018
Closest in time.
Improving language understanding with unsupervised learning
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Closest in time.
U-net: Machine reading comprehension with unanswerable questions
Fu Sun, Linyang Li, Xipeng Qiu, and Yang Liu. 2018 · 2018
Closest in time.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018a · 2018
Closest in time.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. 2018 · 2018
Closest in time.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman. 2018 · 2018
Closest in time.
QANet: Combining local convolution with global self-attention for reading comprehension
Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V Le. 2018 · 2018
Closest in time.
Swag: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Closest in time.