Fetching the paper…
Reading the bibliography…
Pre-trained language models have been shown to encode linguistic structures, e.g.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
An efficient recognition and syntax-analysis algorithm for context-free languages
Tadao Kasami. 1966 · 1966
Earlier work this paper cites.
Trainable grammars for speech recognition
James K Baker. 1979 · 1979
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitch Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Parsing algorithms and metrics
Joshua Goodman. 1996 · 1996
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal. 2020a · 2009
Earlier work this paper cites.
Probing pretrained language models for lexical semantics
Ivan Vulić, Edoardo Maria Ponti, Robert Litschko, Goran Glavaš, and Anna Korhonen. 2020 · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Earlier work this paper cites.
Spectral learning of latent-variable pcfgs
Shay B Cohen, Karl Stratos, Michael Collins, Dean Foster, and Lyle Ungar. 2012 · 2012
Earlier work this paper cites.
Spectral learning of latent-variable pcfgs: Algorithms and sample complexity
Shay B Cohen, Karl Stratos, Michael Collins, Dean P Foster, and Lyle Ungar. 2014 · 2014
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Visualisation and’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
What does bert learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Cited alongside, same era.
Visualizing and measuring the geometry of bert
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019 · 2019
Cited alongside, same era.
Are pre-trained language models aware of phrases? simple but strong baselines for grammar induction
Taeuk Kim, Jihun Choi, Daniel Edmiston, and Sang-goo Lee. 2020 · 2020
Cited alongside, same era.
An empirical comparison of unsupervised constituency parsing methods
Jun Li, Yifan Cao, Jiong Cai, Yong Jiang, and Kewei Tu. 2020 · 2020
Do syntactic probes probe syntax? experiments with jabberwocky probing
Rowan Hall Maudslay and Ryan Cotterell. 2021 · 2021
Later among the works it cites.
Spectral learning of latent-variable pcfgs: High-performance implementation
Haoran Peng. 2021 · 2021
Later among the works it cites.
Attention is turing-complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic. 2021 · 2021
Later among the works it cites.
Syntax-enhanced pre-trained model
Zenan Xu, Daya Guo, Duyu Tang, Qinliang Su, Linjun Shou, Ming Gong, Wanjun Zhong, Xiaojun Quan, Daxin Jiang, and Nan Duan. 2021 · 2021
Later among the works it cites.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos Papadimitriou, and Karthik Narasimhan. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Christopher D Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy. 2020 · 2020
Cited alongside, same era.
A tale of a probe and a parser
Rowan Hall Maudslay, Josef Valvoda, Tiago Pimentel, Adina Williams, and Ryan Cotterell. 2020 · 2020
Cited alongside, same era.
Probing natural language inference models through semantic fragments
Kyle Richardson, Hai Hu, Lawrence Moss, and Ashish Sabharwal. 2020 · 2020
Cited alongside, same era.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Cited alongside, same era.
Parsing as pretraining
David Vilares, Michalina Strzyz, Anders Søgaard, and Carlos Gómez-Rodríguez. 2020 · 2020
Cited alongside, same era.
Perturbed masking: Parameter-free probing for analyzing and interpreting bert
Zhiyong Wu, Yun Chen, Ben Kao, and Qun Liu. 2020 · 2020
Cited alongside, same era.
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. 2022 · 2022
Later among the works it cites.
Probing for constituency structure in neural language models
David Arps, Younes Samih, Laura Kallmeyer, and Hassan Sajjad. 2022 · 2022
Later among the works it cites.
Probing for predicate argument structures in pretrained language models
Simone Conia and Roberto Navigli. 2022 · 2022
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang. 2022 · 2022
Later among the works it cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2022 · 2022
Later among the works it cites.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A Smith. 2022 · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2022 · 2022
Later among the works it cites.
Statistically meaningful approximation: a case study on approximating turing machines with transformers
Colin Wei, Yining Chen, and Tengyu Ma. 2022 · 2022
Later among the works it cites.
Should you mask 15% in masked language modeling?
Alexander Wettig, Tianyu Gao, Zexuan Zhong, and Danqi Chen. 2022 · 2022
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt. 2023 · 2023
Closest in time.