Fetching the paper…
Reading the bibliography…
Generalization remains a central challenge in machine learning.
Ridge regression: Biased estimation for nonorthogonal problems
Arthur E Hoerl and Robert W Kennard · 1970
Earlier work this paper cites.
I I -Divergence Geometry of Probability Distributions and Minimization Problems
I. Csiszar · 1975
Earlier work this paper cites.
The need for biases in learning generalizations
Tom M Mitchell · 1980
Earlier work this paper cites.
A theory and methodology of inductive learning
Ryszard S Michalski · 1983
Earlier work this paper cites.
Learning internal representations by error propagation, 1985
David E Rumelhart, Geoffrey E Hinton, Ronald J Williams, et al · 1985
Earlier work this paper cites.
What size net gives valid generalization?
Eric Baum and David Haussler · 1988
Earlier work this paper cites.
Quantifying inductive bias: Ai learning algorithms and valiant’s learning framework
David Haussler · 1988
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1994
Earlier work this paper cites.
For valid generalization the size of the weights is more important than the size of the network
Peter Bartlett · 1996
Earlier work this paper cites.
Born again trees
Leo Breiman and Nong Shang · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The lasso method for variable selection in the cox model
R Tibshirani · 1997
Earlier work this paper cites.
Ridge regression learning algorithm in dual variables
Craig Saunders, Alexander Gammerman, and Volodya Vovk · 1998
Earlier work this paper cites.
The human adaptation for culture
Michael Tomasello · 1999
Earlier work this paper cites.
Spontaneous evolution of linguistic structure-an iterated learning model of the emergence of regularity and irregularity
Simon Kirby · 2001
Earlier work this paper cites.
Iterated learning: A framework for the emergence of language
Kenny Smith, Simon Kirby, and Henry Brighton · 2003
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 2004
Earlier work this paper cites.
Learning the kernel function via regularization
Charles A Micchelli, Massimiliano Pontil, and Peter Bartlett · 2005
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Cumulative cultural evolution in the laboratory: An experimental approach to the origins of structure in human language
Simon Kirby, Hannah Cornish, and Kenny Smith · 2008
Earlier work this paper cites.
Convention: A philosophical study
David Lewis · 2008
Earlier work this paper cites.
An introduction to Kolmogorov complexity and its applications , volume 3
Ming Li, Paul Vitányi, et al · 2008
Earlier work this paper cites.
Deep learning with kernel regularization for visual recognition
Kai Yu, Wei Xu, and Yihong Gong · 2008
Earlier work this paper cites.
L2 regularization for learning kernels
Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A Krizhevsky · 2009
Earlier work this paper cites.
On the theory of learnining with privileged information
Dmitry Pechyony and Vladimir Vapnik · 2010
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Adaptive dropout for training deep neural networks
Jimmy Ba and Brendan Frey · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
Iterated learning and the evolution of language
Simon Kirby, Tom Griffiths, and Kenny Smith · 2014
Earlier work this paper cites.
Analyzing noise in autoencoders and deep networks
Ben Poole, Jascha Sohl-Dickstein, and Surya Ganguli · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Unifying distillation and privileged information
David Lopez-Paz, Léon Bottou, Bernhard Schölkopf, and Vladimir Vapnik · 2015
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Russ R Salakhutdinov, and Nati Srebro · 2015
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Survey of dropout methods for deep neural networks
Alex Labach, Hojjat Salehinejad, and Shahrokh Valaee · 2019
Later among the works it cites.
SNIP: SINGLE-SHOT NETWORK PRUNING BASED ON CONNECTION SENSITIVITY
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip Torr · 2019
Later among the works it cites.
Ease-of-teaching and language structure from emergent communication
Fushan Li and Michael Bowling · 2019
Later among the works it cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2015
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando De Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Active long term memory networks
Tommaso Furlanello, Jiaping Zhao, Andrew M Saxe, Laurent Itti, and Bosco S Tjan · 2016
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien · 2017
Cited alongside, same era.
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Later among the works it cites.
Transferring inductive biases through knowledge distillation
Samira Abnar, Mostafa Dehghani, and Willem Zuidema · 2020
Later among the works it cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Compositionality and generalization in emergent languages
Rahma Chaabouni, Eugene Kharitonov, Diane Bouchacourt, Emmanuel Dupoux, and Marco Baroni · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
On convergence and generalization of dropout training
Poorya Mianjy and Raman Arora · 2020
Later among the works it cites.
Distilling knowledge via knowledge review
Pengguang Chen, Shu Liu, Hengshuang Zhao, and Jiaya Jia · 2021
Later among the works it cites.
Espn: Extremely sparse pruned networks
Minsu Cho, Ameya Joshi, and Chinmay Hegde · 2021
Later among the works it cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Later among the works it cites.
Convit: Improving vision transformers with soft convolutional inductive biases
Stéphane d’Ascoli, Hugo Touvron, Matthew L Leavitt, Ari S Morcos, Giulio Biroli, and Levent Sagun · 2021
Later among the works it cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Later among the works it cites.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
The deep bootstrap framework: Good online learners are good offline generalizers
Preetum Nakkiran, Behnam Neyshabur, and Hanie Sedghi · 2021
Later among the works it cites.
Learning student-friendly teacher networks for knowledge distillation
Dae Young Park, Moon-Hyun Cha, Changwook Jeong, Daesin Kim, and Bohyung Han · 2021
Later among the works it cites.
Training data-efficient image transformers via distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou · 2021
Later among the works it cites.
Iterated learning for emergent systematicity in {vqa}
Ankit Vani, Max Schwarzer, Yuchen Lu, Eeshan Dhekane, and Aaron Courville · 2021
Later among the works it cites.
Emergent communication at scale
Rahma Chaabouni, Florian Strub, Florent Altché, Eugene Tarassov, Corentin Tallec, Elnaz Davoodi, Kory Wallace Mathewson, Olivier Tieleman, Angeliki Lazaridou, and Bilal Piot · 2022
Later among the works it cites.
Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms
Shengyi Huang, Rousslan Fernand Julien Dossa, Chang Ye, Jeff Braga, Dipam Chakraborty, Kinal Mehta, and João G.M. Araújo · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al · 2022
Later among the works it cites.
Mixskd: Self-knowledge distillation from mixup for image recognition
Chuanguang Yang, Zhulin An, Helong Zhou, Linhang Cai, Xiang Zhi, Jiwen Wu, Yongjun Xu, and Qian Zhang · 2022
Later among the works it cites.
Decoupled knowledge distillation
Borui Zhao, Quan Cui, Renjie Song, Yiyu Qiu, and Jiajun Liang · 2022
Later among the works it cites.
Self-guidance: Improve deep neural network generalization via knowledge distillation
Zhenzhu Zheng and Xi Peng · 2022
Later among the works it cites.
What makes a language easy to deep-learn?
Lukas Galke, Yoav Ram, and Limor Raviv · 2023
Later among the works it cites.
Micah Goldblum, Marc Finzi, Keefer Rowan, and Andrew Gordon Wilson · 2023
Later among the works it cites.
An observation on generalization
Ilya Sutskever · 2023
Later among the works it cites.
An observation on generalization
Ilya Sutskever · 2023
Later among the works it cites.
Mammoth: Building math generalist models through hybrid instruction tuning
Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen · 2023
Later among the works it cites.
Can llms learn by teaching? a preliminary study
Xuefei Ning, Zifu Wang, Shiyao Li, Zinan Lin, Peiran Yao, Tianyu Fu, Matthew B Blaschko, Guohao Dai, Huazhong Yang, and Yu Wang · 2024
Closest in time.
Improving compositional generalization using iterated learning and simplicial embeddings
Yi Ren, Samuel Lavoie, Michael Galkin, Danica J Sutherland, and Aaron C Courville · 2024
Closest in time.