Fetching the paper…
Reading the bibliography…
Language modeling on large-scale datasets leads to impressive performance gains on various downstream language tasks.
Improved sample complexities for deep networks and robust classification via an all-layer margin
Colin Wei and Tengyu Ma · 1910
Earlier work this paper cites.
Three models for the description of language
Noam Chomsky · 1956
Earlier work this paper cites.
On some families of languages related to the dyck language
Maurice Nivat · 1970
Earlier work this paper cites.
The viterbi algorithm
G David Forney · 1973
Earlier work this paper cites.
The estimation of stochastic context-free grammars using the inside-outside algorithm
Karim Lari and Steve J Young · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Pcfg models of linguistic tree representations
Mark Johnson · 1998
Earlier work this paper cites.
Supervised and unsupervised pcfg adaptation to novel domains
Brian Roark and Michiel Bacchiani · 2003
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Earlier work this paper cites.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2018
Earlier work this paper cites.
Dissecting contextual word embeddings: Architecture and representation
Matthew E Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning · 2019
Earlier work this paper cites.
Do attention heads in bert track syntactic dependencies?
Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R Bowman · 2019
Cited alongside, same era.
Compound probabilistic context-free grammars for grammar induction
Yoon Kim, Chris Dyer, and Alexander M Rush · 2019
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
A mathematical exploration of why language models help solve downstream tasks
Nikunj Saunshi, Sadhika Malladi, and Sanjeev Arora · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
The implicit and explicit regularization effects of dropout
Colin Wei, Sham Kakade, and Tengyu Ma · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
Label noise sgd provably prefers flat global minimizers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Cited alongside, same era.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar · 2019
Cited alongside, same era.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Yu Bai and Jason D Lee · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Cited alongside, same era.
Funnel-transformer: Filtering out sequential redundancy for efficient language processing
Zihang Dai, Guokun Lai, Yiming Yang, and Quoc Le · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Alex Damian, Tengyu Ma, and Jason D Lee · 2021
Later among the works it cites.
Provable guarantees for self-supervised deep learning with spectral contrastive loss
Jeff Z HaoChen, Colin Wei, Adrien Gaidon, and Tengyu Ma · 2021
Later among the works it cites.
Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish · 2021
Later among the works it cites.
How to train bert with an academic budget
Peter Izsak, Moshe Berchansky, and Omer Levy · 2021
Later among the works it cites.
What happens after sgd reaches zero loss?–a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2021
Later among the works it cites.
Pay attention to mlps
Hanxiao Liu, Zihang Dai, David So, and Quoc V Le · 2021
Later among the works it cites.
Meta-learning to improve pre-training
Aniruddh Raghu, Jonathan Lorraine, Simon Kornblith, Matthew McDermott, and David K Duvenaud · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al · 2021
Later among the works it cites.
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2021
Later among the works it cites.
Neural bi-lexicalized pcfg induction
Songlin Yang, Yanpeng Zhao, and Kewei Tu · 2021
Later among the works it cites.
On the inductive bias of masked language modeling: From statistical to syntactic dependencies
Tianyi Zhang and Tatsunori Hashimoto · 2021
Later among the works it cites.
Understanding gradient descent on the edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Closest in time.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Closest in time.
Understanding contrastive learning requires incorporating inductive biases
Nikunj Saunshi, Jordan Ash, Surbhi Goel, Dipendra Misra, Cyril Zhang, Sanjeev Arora, Sham Kakade, and Akshay Krishnamurthy · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Closest in time.