Fetching the paper…
Reading the bibliography…
Training large models from scratch usually costs a substantial amount of resources.
Energy-aware neural architecture optimization with fast splitting steepest descent
Dilin Wang, Meng Li, Lemeng Wu, Vikas Chandra, and Qiang Liu · 1910
Earlier work this paper cites.
Lemeng Wu, Mao Ye, Qi Lei, Jason D. Lee, and Qiang Liu · 2003
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2010
Earlier work this paper cites.
Matrix product operator representations
Bogdan Pirvu, Valentin Murg, J Ignacio Cirac, and Frank Verstraete · 2010
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Net2net: Accelerating learning via knowledge transfer
Tianqi Chen, Ian J. Goodfellow, and Jonathon Shlens · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases
Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and Ronald M. Summers · 2017
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Earlier work this paper cites.
Superneurons: dynamic GPU memory management for training deep neural networks
Linnan Wang, Jinmian Ye, Yiyang Zhao, Wei Wu, Ang Li, Shuaiwen Leon Song, Zenglin Xu, and Tim Kraska · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Efficient training of BERT by progressively stacking
Linyuan Gong, Di He, Zhuohan Li, Tao Qin, Liwei Wang, and Tie-Yan Liu · 2019
Earlier work this paper cites.
T-net: Parametrizing fully convolutional nets with a single high-order tensor
Jean Kossaifi, Adrian Bulat, Georgios Tzimiropoulos, and Maja Pantic · 2019
Cited alongside, same era.
A tensorized transformer for language modeling
Xindian Ma, Peng Zhang, Shuai Zhang, Nan Duan, Yuexian Hou, Ming Zhou, and Dawei Song · 2019
Cited alongside, same era.
Compressing recurrent neural networks with tensor ring for action recognition
Yu Pan, Jing Xu, Maolin Wang, Jinmian Ye, Fei Wang, Kun Bai, and Zenglin Xu · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni · 2019
Cited alongside, same era.
On the transformer growth for progressive BERT training
Xiaotao Gu, Liyuan Liu, Hongkun Yu, Jing Li, Chen Chen, and Jiawei Han · 2021
Later among the works it cites.
Transgan: Two pure transformers can make one strong gan, and that can scale up
Yifan Jiang, Shiyu Chang, and Zhangyang Wang · 2021
Later among the works it cites.
Heuristic rank selection with progressively searching tensor ring network
Nannan Li, Yu Pan, Yaran Chen, Zixiang Ding, Dongbin Zhao, and Zenglin Xu · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Tensor ring restricted boltzmann machines
Maolin Wang, Chenbin Zhang, Yu Pan, Jing Xu, and Zenglin Xu · 2019
Cited alongside, same era.
Splitting steepest descent for growing neural architectures
Lemeng Wu, Dilin Wang, and Qiang Liu · 2019
Cited alongside, same era.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Batch normalization biases residual blocks towards the identity function in deep networks
Soham De and Samuel L. Smith · 2020
Cited alongside, same era.
Low-rank compression of neural nets: Learning the rank of each layer
Yerlan Idelbayev and Miguel Á. Carreira-Perpiñán · 2020
Cited alongside, same era.
Qiyu Wu, Chen Xing, Yatao Li, Guolin Ke, Di He, and Tie-Yan Liu · 2021
Later among the works it cites.
Towards efficient tensor decomposition-based DNN model compression with optimization framework
Miao Yin, Yang Sui, Siyu Liao, and Bo Yuan · 2021
Later among the works it cites.
Zero initialization: Initializing residual networks with only zeros and ones
Jiawei Zhao, Florian Schäfer, and Anima Anandkumar · 2021
Later among the works it cites.
Token merging: Your vit but faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman · 2022
Later among the works it cites.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
Fedpara: Low-rank hadamard product for communication-efficient federated learning
Nam Hyeon-Woo, Moon Ye-Bin, and Tae-Hyun Oh · 2022
Later among the works it cites.
Exploring low rank training of deep neural networks
Siddhartha Rao Kamalakara, Acyr Locatelli, Bharat Venkitesh, Jimmy Ba, Yarin Gal, and Aidan N. Gomez · 2022
Later among the works it cites.
Automated progressive learning for efficient training of vision transformers
Changlin Li, Bohan Zhuang, Guangrun Wang, Xiaodan Liang, Xiaojun Chang, and Yi Yang · 2022
Later among the works it cites.
A unified weight initialization paradigm for tensorial convolutional neural networks
Yu Pan, Zeyong Su, Ao Liu, Jingquan Wang, Nannan Li, and Zenglin Xu · 2022
Later among the works it cites.
Knowledge inheritance for pre-trained language models
Yujia Qin, Yankai Lin, Jing Yi, Jiajie Zhang, Xu Han, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou · 2022
Later among the works it cites.
Staged training for transformer language models
Sheng Shen, Pete Walsh, Kurt Keutzer, Jesse Dodge, Matthew E. Peters, and Iz Beltagy · 2022
Later among the works it cites.
Efficienttrain: Exploring generalized curriculum learning for training visual backbones
Yulin Wang, Yang Yue, Rui Lu, Tianjiao Liu, Zhao Zhong, Shiji Song, and Gao Huang · 2022
Later among the works it cites.
Expression syntax information bottleneck for math word problems
Jing Xiong, Chengming Li, Min Yang, Xiping Hu, and Bin Hu · 2022
Later among the works it cites.
Budgeted training for vision transformer
Zhuofan Xia, Xuran Pan, Xuan Jin, Yuan He, Hui Xue, Shiji Song, and Gao Huang · 2023
Closest in time.
Dq-lore: Dual queries with low rank approximation re-ranking for in-context learning
Jiong Xiong, Zixuan Li, Chuanyang Zheng, Zhijiang Guo, Yichun Yin, Enze Xie, Zhicheng Yang, Qingxing Cao, Haiming Wang, Xiongwei Han, Jing Tang, Chengming Li, and Xiaodan Liang · 2023
Closest in time.
Fedpetuning: When federated learning meets the parameter-efficient tuning methods of pre-trained language models
Zhuo Zhang, Yuanhang Yang, Yong Dai, Qifan Wang, Yue Yu, Lizhen Qu, and Zenglin Xu · 2023
Closest in time.