Fetching the paper…
Reading the bibliography…
Classifiers that are linear in their parameters, and trained by optimizing a convex loss function, have predictable behavior with respect to changes in the training data, initial conditions, and optimization.
Why least squares and maximum entropy? an axiomatic approach to inference for linear inverse problems
Imre Csiszar et al · 1991
Earlier work this paper cites.
A visual vocabulary for flower classification
Maria-Elena Nilsback and Andrew Zisserman · 2006
Earlier work this paper cites.
Online linear regression and its application to model-based reinforcement learning
Alexander Strehl and Michael Littman · 2007
Earlier work this paper cites.
Differential privacy: A survey of results
Cynthia Dwork · 2008
Earlier work this paper cites.
Recognizing indoor scenes
Ariadna Quattoni and Antonio Torralba · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Caltech-UCSD Birds 200
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona · 2010
Earlier work this paper cites.
Novel dataset for fine-grained image categorization
Aditya Khosla, Nityananda Jayadevaprakash, Bangpeng Yao, and Li Fei-Fei · 2011
Earlier work this paper cites.
Cats and dogs
Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
Cross-entropy vs. squared error training: a theoretical and experimental comparison
Pavel Golik, Patrick Doetsch, and Hermann Ney · 2013
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
S. Maji, J. Kannala, E. Rahtu, M. Blaschko, and A. Vedaldi · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Bilinear cnn models for fine-grained visual recognition
Tsung-Yu Lin, Aruni RoyChowdhury, and Subhransu Maji · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Deeply learned face representations are sparse, selective, and robust
Yi Sun, Xiaogang Wang, and Xiaoou Tang · 2015
Cited alongside, same era.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Cited alongside, same era.
Time matters in regularizing deep networks: Weight decay and data augmentation affect early learning dynamics, matter little near convergence
Aditya Golatkar, Alessandro Achille, and Stefano Soatto · 2019
Later among the works it cites.
Certified data removal from machine learning models
Chuan Guo, Tom Goldstein, Awni Hannun, and Laurens van der Maaten · 2019
Later among the works it cites.
The ethical algorithm: The science of socially aware algorithm design
Michael Kearns and Aaron Roth · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
Taylorized training: Towards better approximation of neural network training at finite width
Yu Bai, Ben Krause, Huan Wang, Caiming Xiong, and Richard Socher · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tsung-Yu Lin and Subhransu Maji · 2017
Cited alongside, same era.
Large scale fine-grained categorization and domain-specific transfer learning
Yin Cui, Yang Song, Chen Sun, Andrew Howard, and Serge Belongie · 2018
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Identifying medical diagnoses and treatable diseases by image-based deep learning
Daniel S Kermany, Michael Goldbaum, Wenjia Cai, Carolina CS Valentim, Huiying Liang, Sally L Baxter, Alex McKeown, Ge Yang, Xiaokang Wu, Fangbing Yan, et al · 2018
Cited alongside, same era.
Explicit inductive bias for transfer learning with convolutional networks
Xuhong Li, Yves Grandvalet, and Franck Davoine · 2018
Cited alongside, same era.
Pre-trained convolutional neural networks as feature extractors toward improved malaria parasite detection in thin blood smear images
Sivaramakrishnan Rajaraman, Sameer K Antani, Mahdieh Poostchi, Kamolrat Silamut, Md A Hossain, Richard J Maude, Stefan Jaeger, and George R Thoma · 2018
Cited alongside, same era.
Closest in time.
Deep learning on small datasets without pre-training using cosine loss
Bjorn Barz and Joachim Denzler · 2020
Closest in time.
Influence functions in deep learning are fragile
Samyadeep Basu, Philip Pope, and Soheil Feizi · 2020
Closest in time.
Forgetting outside the box: Scrubbing deep networks of information accessible from input-output observations
Aditya Golatkar, Alessandro Achille, and Stefano Soatto · 2020
Closest in time.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Like Hui and Mikhail Belkin · 2020
Closest in time.
What’s in a loss function for image classification?
Simon Kornblith, Honglak Lee, Ting Chen, and Mohammad Norouzi · 2020
Closest in time.
Rethinking the hyperparameters for fine-tuning
Hao Li, Pratik Chaudhari, Hao Yang, Michael Lam, Avinash Ravichandran, Rahul Bhotika, and Stefano Soatto · 2020
Closest in time.
Gradients as features for deep representation learning
Fangzhou Mu, Yingyu Liang, and Yin Li · 2020
Closest in time.
Leep: A new measure to evaluate transferability of learned representations
Cuong V Nguyen, Tal Hassner, Cedric Archambeau, and Matthias Seeger · 2020
Closest in time.
Towards backward-compatible representation learning
Yantao Shen, Yuanjun Xiong, Wei Xia, and Stefano Soatto · 2020
Closest in time.
Estimating informativeness of samples with smooth unique information
Anonymous · 2021
Closest in time.