Bert rediscovers the classical nlp pipeline
Original
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 1905
Earlier work this paper cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Original
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al · 1905
Earlier work this paper cites.
An introduction to hidden markov models
Lawrence Rabiner and Biinghwang Juang · 1986
Earlier work this paper cites.
Robust part-of-speech tagging using a hidden markov model
Julian Kupiec · 1992
Earlier work this paper cites.
Exploiting cloze questions for few shot text classification and natural language inference
Original
Timo Schick and Hinrich Schütze · 2001
Earlier work this paper cites.
Hidden topic markov models
Amit Gruber, Yair Weiss, and Michal Rosen-Zvi · 2007
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
It’s not just size that matters: Small language models are also few-shot learners
Original
Timo Schick and Hinrich Schütze · 2009
Earlier work this paper cites.
Memory networks
Original
Jason Weston, Sumit Chopra, and Antoine Bordes · 2014
Earlier work this paper cites.
End-to-end memory networks
Original
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus · 2015
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Original
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema · 2018
Earlier work this paper cites.
Deep contextualized word representations
Original
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
A theoretical analysis of contrastive unsupervised representation learning
Sanjeev Arora, Hrishikesh Khandeparkar, Mikhail Khodak, Orestis Plevrakis, and Nikunj Saunshi · 2019
Earlier work this paper cites.