Fetching the paper…
Reading the bibliography…
Training deep networks requires various design decisions regarding for instance their architecture, data augmentation, or optimization.
Convolutional Networks for Images, Speech, and Time-Series , page 255–258
Yann LeCun and Yoshua Bengio · 1995
Earlier work this paper cites.
Ensemble methods in machine learning
Thomas G. Dietterich · 2000
Earlier work this paper cites.
Model compression
Cristian Bucila, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Caltech-256 object category dataset
Gregory Griffin, Alex Holub, and Pietro Perona · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Caltech-ucsd birds 200
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2016
Earlier work this paper cites.
Learning without forgetting
Zhizhong Li and Derek Hoiem · 2016
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, G. Sperl, and Christoph H. Lampert · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Earlier work this paper cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis · 2017
Earlier work this paper cites.
Continual learning through synaptic intelligence
Friedemann Zenke, Ben Poole, and Surya Ganguli · 2017
Earlier work this paper cites.
End-to-end incremental learning
Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Cordelia Schmid, and Karteek Alahari · 2018
Earlier work this paper cites.
Bag of tricks for image classification with convolutional neural networks
Tong He, Zhi Zhang, Hang Zhang, Zhongyue Zhang, Junyuan Xie, and Mu Li · 2018
Earlier work this paper cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Arun Mallya and Svetlana Lazebnik · 2018
Earlier work this paper cites.
Piggyback: Adapting a single network to multiple tasks by learning to mask weights
Arun Mallya, Dillon Davis, and Svetlana Lazebnik · 2018
Earlier work this paper cites.
Progress and compress: A scalable framework for continual learning
Jonathan Schwarz, Wojciech Czarnecki, Jelena Luketina, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell · 2018
Earlier work this paper cites.
Deep layer aggregation
Fisher Yu, Dequan Wang, Evan Shelhamer, and Trevor Darrell · 2018
Cited alongside, same era.
Gradient based sample selection for online continual learning
Rahaf Aljundi, Min Lin, Baptiste Goujaud, and Yoshua Bengio · 2019
Cited alongside, same era.
Efficient lifelong learning with a-gem
Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed Elhoseiny · 2019
Cited alongside, same era.
Knowledge flow: Improve upon your teachers
Iou-Jen Liu, Jian Peng, and Alexander G. Schwing · 2019
Cited alongside, same era.
Knowledge amalgamation from heterogeneous networks by common feature learning
Sihui Luo, Xinchao Wang, Gongfan Fang, Yao Hu, Dapeng Tao, and Mingli Song · 2019
Cited alongside, same era.
Moment matching for multi-source domain adaptation
Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang · 2019
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Simultaneous similarity-based self-distillation for deep metric learning
Karsten Roth, Timo Milbich, Bjorn Ommer, Joseph Paul Cohen, and Marzyeh Ghassemi · 2021
Later among the works it cites.
Descending through a crowded valley - benchmarking deep learning optimizers
Robin M Schmidt, Frank Schneider, and Philipp Hennig · 2021
Later among the works it cites.
Dibs: Diversity inducing information bottleneck in model ensembles
Samarth Sinha, Homanga Bharadhwaj, Anirudh Goyal, Hugo Larochelle, Animesh Garg, and Florian Shkurti · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mixconv: Mixed depthwise convolutional kernels
Mingxing Tan and Quoc Chen · 2019
Cited alongside, same era.
Pytorch image models, 2019
Ross Wightman · 2019
Cited alongside, same era.
Dark experience for general continual learning: a strong, simple baseline
Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, and Simone Calderara · 2020
Cited alongside, same era.
Adaptive multi-teacher multi-level knowledge distillation
Yuang Liu, W. Zhang, and Jun Wang · 2020
Cited alongside, same era.
What is being transferred in transfer learning?
Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang · 2020
Cited alongside, same era.
Gdumb: A simple approach that questions our progress in continual learning
Ameya Prabhu, Philip H. S. Torr, and Puneet Kumar Dokania · 2020
Cited alongside, same era.
MLP-mixer: An all-MLP architecture for vision
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Peter Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Resmlp: Feedforward networks for image classification with data-efficient training
Hugo Touvron, Piotr Bojanowski, Mathilde Caron, Matthieu Cord, Alaaeldin El-Nouby, Edouard Grave, Gautier Izacard, Armand Joulin, Gabriel Synnaeve, Jakob Verbeek, and Hervé Jégou · 2021
Later among the works it cites.
One teacher is enough? pre-trained language model distillation from multiple teachers
Chuhan Wu, Fangzhao Wu, and Yongfeng Huang · 2021
Later among the works it cites.
Self-supervised learning with swin transformers
ZeLun Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2021
Later among the works it cites.
Co-scale conv-attentional image transformers
Weijian Xu, Yifan Xu, Tyler Chang, and Zhuowen Tu · 2021
Later among the works it cites.
Volo: Vision outlooker for visual recognition
Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Knowledge distillation: A good teacher is patient and consistent
Lucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov · 2022
Later among the works it cites.
Class-incremental learning via knowledge amalgamation
Marcus Vinícius de Carvalho, Mahardhika Pratama, Jie Zhang, and Yajuan San · 2022
Later among the works it cites.
No one representation to rule them all: Overlapping features of training methods
Raphael Gontijo-Lopes, Yann Dauphin, and Ekin Dogus Cubuk · 2022
Later among the works it cites.
FFCV: Accelerating training by removing data bottlenecks
Guillaume Leclerc, Andrew Ilyas, Logan Engstrom, Sung Min Park, Hadi Salman, and Aleksander Madry · 2022
Later among the works it cites.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Later among the works it cites.
No one representation to rule them all: Overlapping features of training methods
Raphael Gontijo Lopes, Yann Dauphin, and Ekin Dogus Cubuk · 2022
Later among the works it cites.
Integrating language guidance into vision-based deep metric learning
Karsten Roth, Oriol Vinyals, and Zeynep Akata · 2022
Later among the works it cites.
Momentum-based weight interpolation of strong zero-shot models for continual learning
Zafir Stojanovski, Karsten Roth, and Zeynep Akata · 2022
Later among the works it cites.
On the importance of hyperparameters and data augmentation for self-supervised learning
Diane Wagner, Fabio Ferreira, Danny Stoll, Robin Tibor Schirrmeister, Samuel Müller, and Frank Hutter · 2022
Later among the works it cites.
Multi-teacher knowledge distillation for incremental implicitly-refined classification
Longhui Yu, Zhenyu Weng, Yuqing Wang, and Yuesheng Zhu · 2022
Later among the works it cites.
A cookbook of self-supervised learning
Randall Balestriero, Mark Ibrahim, Vlad Sobal, Ari Morcos, Shashank Shekhar, Tom Goldstein, Florian Bordes, Adrien Bardes, Gregoire Mialon, Yuandong Tian, Avi Schwarzschild, Andrew Gordon Wilson, Jonas Geiping, Quentin Garrido, Pierre Fernandez, Amir Bar, Hamed Pirsiavash, Yann LeCun, and Micah Goldblum · 2023
Closest in time.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision, 2023
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeff Wu · 2023
Closest in time.
Disentanglement of correlated factors via hausdorff factorized support
Karsten Roth, Mark Ibrahim, Zeynep Akata, Pascal Vincent, and Diane Bouchacourt · 2023
Closest in time.
Can cnns be more robust than transformers?
Zeyu Wang, Yutong Bai, Yuyin Zhou, and Cihang Xie · 2023
Closest in time.