Fetching the paper…
Reading the bibliography…
It is widely believed that deep neural networks contain layer specialization, wherein neural networks extract hierarchical features representing edges and patterns in shallow layers and complete objects in deeper layers.
The principled design of large-scale recursive neural network architectures–dag-rnns and the protein structure prediction problem
Pierre Baldi and Gianluca Pollastri · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Understanding deep architectures using a recursive convolutional network
David Eigen, Jason Rolfe, Rob Fergus, and Yann LeCun · 2013
Earlier work this paper cites.
Recurrent convolutional neural networks for scene labeling
Pedro Pinheiro and Ronan Collobert · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Recurrent convolutional neural networks for text classification
Siwei Lai, Liheng Xu, Kang Liu, and Jun Zhao · 2015
Earlier work this paper cites.
Recurrent convolutional neural network for object recognition
Ming Liang and Xiaolin Hu · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Understanding neural networks through deep visualization
Jason Yosinski, Jeff Clune, Anh Nguyen, Thomas Fuchs, and Hod Lipson · 2015
Earlier work this paper cites.
Highway and residual networks learn unrolled iterative estimation
Klaus Greff, Rupesh K Srivastava, and Jürgen Schmidhuber · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Bridging the gaps between residual learning, recurrent neural networks and visual cortex
Qianli Liao and Tomaso Poggio · 2016
Cited alongside, same era.
Sharesnet: reducing residual network parameter number by sharing weights
Alexandre Boulch · 2017
Cited alongside, same era.
Emnist: an extension of mnist to handwritten letters, 2017
Gregory Cohen, Saeed Afshar, Jonathan Tapson, and André van Schaik · 2017
Cited alongside, same era.
Making a maze, Apr 2017
Christian Hill · 2017
Cited alongside, same era.
Deep equilibrium models
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2019
Later among the works it cites.
An ensemble of lstm neural networks for high-frequency stock market classification
Svetlana Borovkova and Ioannis Tsiamas · 2019
Later among the works it cites.
Shallow-deep networks: Understanding and mitigating network overthinking
Yigitcan Kaya, Sanghyun Hong, and Tudor Dumitras · 2019
Later among the works it cites.
Similarity of neural network representations revisited, 2019
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Residual connections encourage iterative inference
Stanisław Jastrzebski, Devansh Arpit, Nicolas Ballas, Vikas Verma, Tong Che, and Yoshua Bengio · 2017
Cited alongside, same era.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Cited alongside, same era.
Md Zahangir Alom, Mahmudul Hasan, Chris Yakopcic, Tarek M Taha, and Vijayan K Asari · 2018
Cited alongside, same era.
Trellis networks for sequence modeling
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2018
Cited alongside, same era.
Neural ordinary differential equations
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud · 2018
Cited alongside, same era.
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel · 2018
Cited alongside, same era.
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Later among the works it cites.
Adversarial attacks on machine learning systems for high-frequency trading, 2020
Micah Goldblum, Avi Schwarzschild, Ankit B. Patel, and Tom Goldstein · 2020
Later among the works it cites.
Understanding generalization through visualizations
W Ronny Huang, Zeyad Emam, Micah Goldblum, Liam Fowl, Justin K Terry, Furong Huang, and Tom Goldstein · 2020
Later among the works it cites.
Dreaming to distill: Data-free knowledge transfer via deepinversion
Hongxu Yin, Pavlo Molchanov, Jose M Alvarez, Zhizhong Li, Arun Mallya, Derek Hoiem, Niraj K Jha, and Jan Kautz · 2020
Later among the works it cites.
High-performance large-scale image recognition without normalization
Andrew Brock, Soham De, Samuel L Smith, and Karen Simonyan · 2021
Closest in time.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira · 2021
Closest in time.