Megatron-lm: Training multi-billion parameter language models using model parallelism
Original
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Later among the works it cites.
Lookahead optimizer: k steps forward, 1 step back
Michael Zhang, James Lucas, Jimmy Ba, and Geoffrey E Hinton · 2019
Later among the works it cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
Rigging the lottery: Making all tickets winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen · 2020
Later among the works it cites.
Top-kast: Top-k always sparse training
Siddhant Jayakumar, Razvan Pascanu, Jack Rae, Simon Osindero, and Erich Elsen · 2020
Later among the works it cites.
What is being transferred in transfer learning?
Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang · 2020
Later among the works it cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
Original
Hidenori Tanaka, Daniel Kunin, Daniel LK Yamins, and Surya Ganguli · 2020
Later among the works it cites.
Batchensemble: an alternative approach to efficient ensemble and lifelong learning
Original
Yeming Wen, Dustin Tran, and Jimmy Ba · 2020
Later among the works it cites.
Towards structured dynamic sparse pre-training of bert
Original
Anastasia Dietrich, Frithjof Gressmann, Douglas Orr, Ivan Chelombiev, Daniel Justus, and Carlo Luschi · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Original
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Original
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Learning neural network subspaces
Mitchell Wortsman, Maxwell C Horton, Carlos Guestrin, Ali Farhadi, and Mohammad Rastegari · 2021
Later among the works it cites.
Mest: Accurate and fast memory-economic sparse training framework on the edge
Geng Yuan, Xiaolong Ma, Wei Niu, Zhengang Li, Zhenglun Kong, Ning Liu, Yifan Gong, Zheng Zhan, Chaoyang He, Qing Jin, et al · 2021
Later among the works it cites.
On heterogeneously distributed data, sparsity matters
Tiansheng Huang, Shiwei Liu, L Shen, Fengxiang He, Weiwei Lin, and Dacheng Tao · 2022
Closest in time.
The unreasonable effectiveness of random pruning: Return of the most naive baseline for sparse training
Original
Shiwei Liu, Tianlong Chen, Xiaohan Chen, Li Shen, Decebal Constantin Mocanu, Zhangyang Wang, and Mykola Pechenizkiy · 2022
Closest in time.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Original
Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al · 2022
Closest in time.