Fetching the paper…
Reading the bibliography…
Recent studies have shown that language models pretrained and/or fine-tuned on randomly permuted sentences exhibit competitive performance on GLUE, putting into question the importance of word order information.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 1905
Earlier work this paper cites.
Open sesame: Getting inside bert’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 1906
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
On the Relationship between Self-Attention and Convolutional Layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. 2020 · 1911
Earlier work this paper cites.
The cognitive basis for linguistic structures
Thomas G Bever. 1970 · 1970
Earlier work this paper cites.
Psychological scaling of adjective orders
Joseph H Danks and Sam Glucksberg. 1971 · 1971
Earlier work this paper cites.
A theory of reading: From eye fixations to comprehension
Marcel A Just and Patricia A Carpenter. 1980 · 1980
Earlier work this paper cites.
Language universals and linguistic typology: Syntax and morphology
Bernard Comrie. 1989 · 1989
Earlier work this paper cites.
Training with noise is equivalent to tikhonov regularization
Chris M. Bishop. 1995 · 1995
Earlier work this paper cites.
Auditory language comprehension: an event-related fmri study on the processing of syntactic and lexical information
Angela D Friederici, Martin Meyer, and D Yves Von Cramon. 2000 · 2000
Earlier work this paper cites.
Syntactic parsing preferences and their on-line revisions: A spatio-temporal analysis of event-related brain potentials
Angela D Friederici, Axel Mecklinger, Kevin M Spencer, Karsten Steinhauer, and Emanuel Donchin. 2001 · 2001
Earlier work this paper cites.
Good-enough representations in language comprehension
Fernanda Ferreira, Karl GD Bailey, and Vittoria Ferraro. 2002 · 2002
Earlier work this paper cites.
A theory of usable information under computational constraints
Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon. 2020 · 2002
Earlier work this paper cites.
Word length, sentence length and frequency – zipf revisited
Bengt Sigurd, Mats Eeg-Olofsson, and Joost Van Weijer. 2004 · 2004
Earlier work this paper cites.
An fmri study of canonical and noncanonical word order in german
Jörg Bahlmann, Antoni Rodriguez-Fornells, Michael Rotte, and Thomas F Münte. 2007 · 2007
Earlier work this paper cites.
Explicit Regularisation in Gaussian Noise Injections
Alexander Camuto, Matthew Willetts, Umut Şimşekli, Stephen Roberts, and Chris Holmes. 2021 · 2007
Earlier work this paper cites.
Topographic mapping of a hierarchy of temporal receptive windows using a narrated story
Yulia Lerner, Christopher J Honey, Lauren J Silbert, and Uri Hasson. 2011 · 2011
Earlier work this paper cites.
Cortical representation of the constituent structure of sentences
Christophe Pallier, Anne-Dominique Devauchelle, and Stanislas Dehaene. 2011 · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011 · 2011
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Thang M Pham, Trung Bui, Long Mai, and Anh Nguyen. 2020 · 2012
Earlier work this paper cites.
Rational integration of noisy evidence and prior semantic expectations in sentence interpretation
Edward Gibson, Leon Bergen, and Steven T Piantadosi. 2013 · 2013
Cited alongside, same era.
A gold standard dependency corpus for english
Natalia Silveira, Timothy Dozat, Marie-Catherine De Marneffe, Samuel R Bowman, Miriam Connor, John Bauer, and Christopher D Manning. 2014 · 2014
Cited alongside, same era.
Trends in syntactic parsing: Anticipation, bayesian estimation, and good-enough parsing
Matthew J Traxler. 2014 · 2014
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Cortical tracking of hierarchical linguistic structures in connected speech
Nai Ding, Lucia Melloni, Hang Zhang, Xing Tian, and David Poeppel. 2016 · 2016
Cited alongside, same era.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Later among the works it cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
SuperGLUE: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural correlate of the construction of sentence meaning
Evelina Fedorenko, Terri L Scott, Peter Brunner, William G Coon, Brianna Pritchett, Gerwin Schalk, and Nancy Kanwisher. 2016 · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
Ankur P Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Cited alongside, same era.
Stanford’s graph-based neural dependency parser at the CoNLL 2017 shared task
Timothy Dozat, Peng Qi, and Christopher D. Manning. 2017 · 2017
Cited alongside, same era.
How will text size influence the length of its linguistic constituents?
Huiyuan Jin and Haitao Liu. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
The gum corpus: Creating multilayer resources in the classroom
Amir Zeldes. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Yuan Zhang, Jason Baldridge, and Luheng He. 2019 · 2019
Later among the works it cites.
Composition is the core driver of the language-selective network
Francis Mollica, Matthew Siegelman, Evgeniia Diachek, Steven T Piantadosi, Zachary Mineroff, Richard Futrell, Hope Kean, Peng Qian, and Evelina Fedorenko. 2020 · 2020
Later among the works it cites.
Pareto probing: Trading off accuracy for complexity
Tiago Pimentel, Naomi Saphra, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, and Adina Williams. 2020 · 2020
Later among the works it cites.
What do position embeddings learn? an empirical study of pre-trained language model positional encoding
Yu-An Wang and Yun-Nung Chen. 2020 · 2020
Later among the works it cites.
Matteo Alleman, Jonathan Mamou, Miguel A Del Rio, Hanlin Tang, Yoon Kim, and SueYeon Chung. 2021 · 2021
Later among the works it cites.
Demystifying neural language models’ insensitivity to word-order
Louis Clouatre, Prasanna Parthasarathi, Amal Zouaq, and Sarath Chandar. 2021 · 2021
Later among the works it cites.
Bert & family eat word salad: Experiments with text understanding
Ashim Gupta, Giorgi Kvernadze, and Vivek Srikumar. 2021 · 2021
Later among the works it cites.
Schrödinger’s Tree – On Syntax and Neural Language Models
Artur Kulmizev and Joakim Nivre. 2021 · 2021
Later among the works it cites.
What context features can transformer language models use?
Joe O’Connor and Jacob Andreas. 2021 · 2021
Later among the works it cites.
Deep subjecthood: Higher-order grammatical features in multilingual bert
Isabel Papadimitriou, Ethan A Chi, Richard Futrell, and Kyle Mahowald. 2021 · 2021
Later among the works it cites.
Attention can reflect syntactic structure (if you let it)
Vinit Ravishankar, Artur Kulmizev, Mostafa Abdou, Anders Søgaard, and Joakim Nivre. 2021 · 2021
Later among the works it cites.
The impact of positional encodings on multilingual compression
Vinit Ravishankar and Anders Søgaard. 2021 · 2021
Later among the works it cites.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
ON POSITION EMBEDDINGS IN BERT
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Hao Yang, Qun Liu, and Jakob Grue Simonsen. 2021 · 2021
Later among the works it cites.