On the computational power of neural nets
Hava T. Siegelmann and Eduardo D. Sontag · 1992
Earlier work this paper cites.
Fully dynamic transitive closure: breaking through the o (n/sup 2/) barrier
Camil Demetrescu and Giuseppe F Italiano · 2000
Earlier work this paper cites.
All pairs shortest paths using bridging sets and rectangular matrix multiplication
Uri Zwick · 2002
Earlier work this paper cites.
Dynamic transitive closure via dynamic matrix inverse
Piotr Sankowski · 2004
Earlier work this paper cites.
Subquadratic algorithm for dynamic shortest distances
Piotr Sankowski · 2005
Earlier work this paper cites.
Multiplying matrices faster than coppersmith-winograd
Virginia Vassilevska Williams · 2012
Earlier work this paper cites.
Powers of tensors and fast matrix multiplication
François Le Gall · 2014
Earlier work this paper cites.
Unifying and strengthening hardness for dynamic problems via the online matrix-vector multiplication conjecture
Monika Henzinger, Sebastian Krinninger, Danupon Nanongkai, and Thatchaphol Saranurak · 2015
Earlier work this paper cites.
On the fine-grained complexity of empirical risk minimization: Kernel methods and neural networks
Arturs Backurs, Piotr Indyk, and Ludwig Schmidt · 2017
Earlier work this paper cites.
Faster online matrix-vector multiplication
Kasper Green Larsen and Ryan Williams · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Tight cell probe bounds for succinct boolean matrix-vector multiplication
Diptarka Chakraborty, Lior Kamma, and Kasper Green Larsen · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
A matrix expander chernoff bound
Ankit Garg, Yin Tat Lee, Zhao Song, and Nikhil Srivastava · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Dynamic approximate shortest paths and beyond: Subquadratic and worst-case update time
Jan van den Brand and Danupon Nanongkai · 2019
Earlier work this paper cites.
Dynamic matrix inverse: Improved algorithms and matching conditional lower bounds
Jan van den Brand, Danupon Nanongkai, and Thatchaphol Saranurak · 2019
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 2019
Earlier work this paper cites.
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning · 2019
Earlier work this paper cites.