Fetching the paper…
Reading the bibliography…
While vector-based language representations from pretrained language models have set a new standard for many NLP tasks, there is not yet a complete accounting of their inner workings.
Assessing bert’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019a · 1903
Earlier work this paper cites.
Hierarchical representation in neural language models: Suppression and recovery of expectations
Ethan Wilcox, Roger Levy, and Richard Futrell. 2019 · 1906
Earlier work this paper cites.
Distributed representations, simple recurrent networks, and grammatical structure
Jeffrey L Elman. 1991 · 1991
Earlier work this paper cites.
The penn treebank: Annotating predicate argument structure
Mitchell Marcus, Grace Kim, Mary Ann Marcinkiewicz, Robert MacIntyre, Ann Bies, Mark Ferguson, Karen Katz, and Britta Schasberger. 1994 · 1994
Earlier work this paper cites.
Econometrics
Bruce Hansen. 2000 · 2000
Earlier work this paper cites.
Emergence of separable manifolds in deep language representations
Jonathan Mamou, Hang Le, Miguel A Del Rio, Cory Stephenson, Hanlin Tang, Yoon Kim, and SueYeon Chung. 2020 · 2006
Earlier work this paper cites.
Untangling invariant object recognition
James J DiCarlo and David D Cox. 2007 · 2007
Earlier work this paper cites.
The discovery of structural form
Charles Kemp and Joshua B Tenenbaum. 2008 · 2008
Earlier work this paper cites.
Mechanisms of face perception
Doris Y Tsao and Margaret S Livingstone. 2008 · 2008
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Learning hierarchical categories in deep neural networks
Andrew M Saxe, James L McClellans, and Surya Ganguli. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015 · 2015
Earlier work this paper cites.
Cortical tracking of hierarchical linguistic structures in connected speech
Nai Ding, Lucia Melloni, Hang Zhang, Xing Tian, and David Poeppel. 2016 · 2016
Earlier work this paper cites.
Learning distributed representations of sentences from unlabelled data
Felix Hill, Kyunghyun Cho, and Anna Korhonen. 2016 · 2016
Cited alongside, same era.
Toward the neural implementation of structure learning
D Gowanlock R Tervo, Joshua B Tenenbaum, and Samuel J Gershman. 2016 · 2016
Cited alongside, same era.
Analyzing hidden representations in end-to-end automatic speech recognition systems
Yonatan Belinkov and James Glass. 2017 · 2017
Cited alongside, same era.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2017
Cited alongside, same era.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein. 2017 · 2017
Cited alongside, same era.
The hippocampus as a predictive map
Kimberly L Stachenfeld, Matthew M Botvinick, and Samuel J Gershman. 2017 · 2017
Neural language models as psycholinguistic subjects: Representations of syntactic state
Richard Futrell, Ethan Wilcox, Takashi Morita, Peng Qian, Miguel Ballesteros, and Roger Levy. 2019 · 2019
Later among the works it cites.
Visualizing the phate of neural networks
Scott Gigante, Adam S Charles, Smita Krishnaswamy, and Gal Mishne. 2019 · 2019
Later among the works it cites.
Designing and Interpreting Probes with Control Tasks
John Hewitt and Percy Liang. 2019 · 2019
Later among the works it cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning. 2019 · 2019
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unsupervised neural machine translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018 · 2018
Cited alongside, same era.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Cited alongside, same era.
Classification and geometry of general perceptual manifolds
SueYeon Chung, Daniel D Lee, and Haim Sompolinsky. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
RNNs as psycholinguistic subjects: Syntactic state and grammatical dependency
Richard Futrell, Ethan Wilcox, Takashi Morit, and Roger Levy. 2018 · 2018
Cited alongside, same era.
Unsupervised machine translation using monolingual corpora only
Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018 · 2018
Cited alongside, same era.
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019 · 2019
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
Hierarchical reasoning by neural circuits in the frontal cortex
Morteza Sarafyazd and Mehrdad Jazayeri. 2019 · 2019
Later among the works it cites.
BERT Rediscovers the Classical NLP Pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
Good-Enough Compositional Data Augmentation
Jacob Andreas. 2020 · 2020
Later among the works it cites.
What bert is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2020 · 2020
Later among the works it cites.
Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar Induction
Taeuk Kim, Jihun Choi, Daniel Edmiston, and Sang goo Lee. 2020 · 2020
Later among the works it cites.
Composition is the core driver of the language-selective network
Francis Mollica, Matthew Siegelman, Evgeniia Diachek, Steven T Piantadosi, Zachary Mineroff, Richard Futrell, Hope Kean, Peng Qian, and Evelina Fedorenko. 2020 · 2020
Later among the works it cites.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021 · 2021
Closest in time.