A winning hand: Compressing deep networks can improve out-of-distribution robustness
James Diffenderfer, Brian Bartoldson, Shreya Chaganti, Jize Zhang, and Bhavya Kailkhura · 2021
Later among the works it cites.
Pruning neural networks at initialization: Why are we missing the mark?
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2021
Later among the works it cites.
M-fac: Efficient matrix-free approximations of second-order information
Elias Frantar, Eldar Kurtic, and Dan Alistarh · 2021
Later among the works it cites.
Towards more effective and economic sparsely-activated model
Original
Hao Jiang, Ke Zhan, Jianwei Qu, Yongkang Wu, Zhaoye Fei, Xinyu Zhang, Lei Chen, Zhicheng Dou, Xipeng Qiu, Zikai Guo, et al · 2021
Later among the works it cites.
Block pruning for faster transformers
Original
François Lagunas, Ella Charlaix, Victor Sanh, and Alexander M Rush · 2021
Later among the works it cites.
A diverse corpus for evaluating and developing english math word problem solvers
Original
Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su · 2021
Later among the works it cites.
Training adversarially robust sparse networks via bayesian connectivity sampling
Ozan Özdenizci and Robert Legenstein · 2021
Later among the works it cites.
Are nlp models really able to solve simple math word problems?
Original
Arkil Patel, Satwik Bhattamishra, and Navin Goyal · 2021
Later among the works it cites.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al · 2021
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2021
Later among the works it cites.
Powerpropagation: A sparsity inducing weight reparameterisation
Jonathan Schwarz, Siddhant Jayakumar, Razvan Pascanu, Peter E Latham, and Yee Teh · 2021
Later among the works it cites.
Rethinking network pruning–under the pre-train and fine-tune paradigm
Original
Dongkuan Xu, Ian EH Yen, Jinxi Zhao, and Zhibin Xiao · 2021
Later among the works it cites.
Prune once for all: Sparse pre-trained language models
Original
Ofir Zafrir, Ariel Larey, Guy Boudoukh, Haihao Shen, and Moshe Wasserblat · 2021
Later among the works it cites.
Can subnetwork structure be the key to out-of-distribution generalization?
Dinghuai Zhang, Kartik Ahuja, Yilun Xu, Yisen Wang, and Aaron Courville · 2021
Later among the works it cites.
Sparsity winning twice: Better robust generalization from more efficient training
Tianlong Chen, Zhenyu Zhang, Santosh Balachandra, Haoyu Ma, Zehao Wang, Zhangyang Wang, et al · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Original
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Learning inverse folding from millions of predicted structures
Chloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexander Rives · 2022
Later among the works it cites.
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
Original
Eldar Kurtic, Daniel Campos, Tuan Nguyen, Elias Frantar, Mark Kurtz, Benjamin Fineran, Michael Goin, and Dan Alistarh · 2022
Later among the works it cites.
The unreasonable effectiveness of random pruning: Return of the most naive baseline for sparse training
Shiwei Liu, Tianlong Chen, Xiaohan Chen, Li Shen, Decebal Constantin Mocanu, Zhangyang Wang, and Mykola Pechenizkiy · 2022
Later among the works it cites.
A kernel-based view of language model fine-tuning
Original
Sadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen, and Sanjeev Arora · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Original
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Original
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Original
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al · 2022
Later among the works it cites.
Lottery pools: Winning more by interpolating tickets without increasing training or inference cost
Original
Lu Yin, Shiwei Liu, Fang Meng, Tianjin Huang, Vlado Menkovski, and Mykola Pechenizkiy · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Original
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Later among the works it cites.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Later among the works it cites.
Platon: Pruning large transformer models with upper confidence bound of weight importance
Qingru Zhang, Simiao Zuo, Chen Liang, Alexander Bukharin, Pengcheng He, Weizhu Chen, and Tuo Zhao · 2022
Later among the works it cites.
Hotprotein: A novel framework for protein thermostability prediction and editing
Tianlong Chen, Chengyue Gong, Daniel Jesus Diaz, Xuxi Chen, Jordan Tyler Wells, qiang liu, Zhangyang Wang, Andrew Ellington, Alex Dimakis, and Adam Klivans · 2023
Closest in time.
Ten lessons we have learned in the new” sparseland”: A short handbook for sparse neural network researchers
Original
Shiwei Liu and Zhangyang Wang · 2023
Closest in time.