Large scale distributed neural network training through online distillation
Rohan Anil, Gabriel Pereyra, Alexandre Tachard Passos, Robert Ormandi, George Dahl, and Geoffrey Hinton · 2018
Later among the works it cites.
Born-again neural networks
Original
Tommaso Furlanello, Zachary Chase Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Later among the works it cites.
On the information bottleneck theory of deep learning
Andrew Michael Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan Daniel Tracey, and David Daniel Cox · 2018
Later among the works it cites.
The importance of being recurrent for modeling hierarchical structure
Ke Tran, Arianna Bisazza, and Christof Monz · 2018
Later among the works it cites.
Blackbox meets blackbox: Representational similarity & stability analysis of neural language models and brains
Samira Abnar, Lisa Beinborn, Rochelle Choenni, and Willem Zuidema · 2019
Later among the works it cites.
Universal transformers
Original
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Later among the works it cites.
Modeling recurrence for transformer
Original
Jie Hao, Xing Wang, Baosong Yang, Longyue Wang, Jinfeng Zhang, and Zhaopeng Tu · 2019
Later among the works it cites.
Scalable syntax-aware language models using knowledge distillation
Adhiguna Kuncoro, Chris Dyer, Laura Rimell, Stephen Clark, and Phil Blunsom · 2019
Later among the works it cites.
Mnist-c: A robustness benchmark for computer vision
Original
Norman Mu and Justin Gilmer · 2019
Later among the works it cites.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 2019
Later among the works it cites.
Language models are few-shot learners
Original
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Closest in time.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Closest in time.
Syntactic structure distillation pretraining for bidirectional encoders
Original
Adhiguna Kuncoro, Lingpeng Kong, Daniel Fried, Dani Yogatama, Laura Rimell, Chris Dyer, and Phil Blunsom · 2020
Closest in time.
Self-distillation amplifies regularization in hilbert space
Original
Hossein Mobahi, Mehrdad Farajtabar, and Peter L Bartlett · 2020
Closest in time.
What is being transferred in transfer learning?
Original
Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang · 2020
Closest in time.
Understanding and improving knowledge distillation
Original
Jiaxi Tang, Rakesh Shivanna, Zhe Zhao, Dong Lin, Anima Singh, Ed H Chi, and Sagar Jain · 2020
Closest in time.