Fetching the paper…
Reading the bibliography…
Deep learning has been successful in automating the design of features in machine learning pipelines.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma · 1901
Earlier work this paper cites.
Knowledge flow: Improve upon your teachers
Iou-Jen Liu, Jian Peng, and Alexander G Schwing · 1904
Earlier work this paper cites.
Why adam beats sgd for attention models
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra · 1912
Earlier work this paper cites.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
The fréchet distance between multivariate normal distributions
DC Dowson and BV Landau · 1982
Earlier work this paper cites.
Metalearning machines learn to learn (1987-)
Jürgen Schmidhuber and AI Blog · 1987
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
Salah El Hihi and Yoshua Bengio · 1996
Earlier work this paper cites.
The architecture of complex weighted networks
Alain Barrat, Marc Barthelemy, Romualdo Pastor-Satorras, and Alessandro Vespignani · 2004
Earlier work this paper cites.
The computational limits of deep learning
Neil C Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F Manso · 2007
Earlier work this paper cites.
Object detection combining recognition and segmentation
Liming Wang, Jianbo Shi, Gang Song, and I-fan Shen · 2007
Earlier work this paper cites.
Exploring network structure, dynamics, and function using networkx
Aric Hagberg, Pieter Swart, and Daniel S Chult · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky et al · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Building lego using deep generative models of graphs
Rylee Thompson, Elahe Ghalebi, Terrance DeVries, and Graham W Taylor · 2012
Earlier work this paper cites.
Predicting parameters in deep learning
Misha Denil, Babak Shakibi, Laurent Dinh, Marc’Aurelio Ranzato, and Nando de Freitas · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Rigid-motion scattering for texture classification
Laurent Sifre and Stéphane Mallat · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Gated graph sequence neural networks
Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard Zemel · 2015
Earlier work this paper cites.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Deep graph kernels
Pinar Yanardag and SVN Vishwanathan · 2015
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V Le · 2016
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Earlier work this paper cites.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
What makes imagenet good for transfer learning?
Minyoung Huh, Pulkit Agrawal, and Alexei A Efros · 2016
Earlier work this paper cites.
Diet networks: thin parameters for fat genomics
Adriana Romero, Pierre Luc Carrier, Akram Erraqabi, Tristan Sylvain, Alex Auvolat, Etienne Dejoie, Marc-André Legault, Marie-Pierre Dubé, Julie G Hussin, and Yoshua Bengio · 2016
Earlier work this paper cites.
Learning feed-forward one-shot learners
Luca Bertinetto, João F Henriques, Jack Valmadre, Philip HS Torr, and Andrea Vedaldi · 2016
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Cited alongside, same era.
Lgm-net: Learning to generate matching networks for few-shot learning
Huaiyu Li, Weiming Dong, Xing Mei, Chongyang Ma, Feiyue Huang, and Bao-Gang Hu · 2019
Later among the works it cites.
Simgnn: A neural network approach to fast graph similarity computation
Yunsheng Bai, Hao Ding, Song Bian, Ting Chen, Yizhou Sun, and Wei Wang · 2019
Later among the works it cites.
The frechet distance of training and test distribution predicts the generalization gap
Julian Zilly, Hannes Zilly, Oliver Richter, Roger Wattenhofer, Andrea Censi, and Emilio Frazzoli · 2019
Later among the works it cites.
Hypergan: A generative model for diverse, performant neural networks
Neale Ratzlaff and Li Fuxin · 2019
Later among the works it cites.
Metainit: Initializing learning by learning to initialize
Yann Dauphin and Samuel S Schoenholz · 2019
Later among the works it cites.
Dag-gnn: Dag structure learning with graph neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Cited alongside, same era.
Impact of training set batch size on the performance of convolutional neural networks for diverse datasets
Pavlo M Radiuk · 2017
Cited alongside, same era.
Learned optimizers that scale and generalize
Olga Wichrowska, Niru Maheswaranathan, Matthew W Hoffman, Sergio Gomez Colmenarejo, Misha Denil, Nando Freitas, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Accelerating neural architecture search using performance prediction
Bowen Baker, Otkrist Gupta, Ramesh Raskar, and Nikhil Naik · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Yue Yu, Jie Chen, Tian Gao, and Mo Yu · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Meta-learning in neural networks: A survey
Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Benchmarking graph neural networks
Vijay Prakash Dwivedi, Chaitanya K Joshi, Thomas Laurent, Yoshua Bengio, and Xavier Bresson · 2020
Later among the works it cites.
Are wider nets better given the same number of parameters?
Anna Golubeva, Behnam Neyshabur, and Guy Gur-Ari · 2020
Later among the works it cites.
The hardware lottery, 2020
Sara Hooker · 2020
Later among the works it cites.
On the bottleneck of graph neural networks and its practical implications
Uri Alon and Eran Yahav · 2020
Later among the works it cites.
Non-local graph neural networks
Meng Liu, Zhengyang Wang, and Shuiwang Ji · 2020
Later among the works it cites.
Geom-gcn: Geometric graph convolutional networks
Hongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei, and Bo Yang · 2020
Later among the works it cites.
Cars: Continuous evolution for efficient neural architecture search
Zhaohui Yang, Yunhe Wang, Xinghao Chen, Boxin Shi, Chao Xu, Chunjing Xu, Qi Tian, and Chang Xu · 2020
Later among the works it cites.
Milenas: Efficient neural architecture search via mixed-level reformulation
Chaoyang He, Haishan Ye, Li Shen, and Tong Zhang · 2020
Later among the works it cites.
Luke Metz, Niru Maheswaranathan, C Daniel Freeman, Ben Poole, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
A survey on graph kernels
Nils M Kriege, Fredrik D Johansson, and Christopher Morris · 2020
Later among the works it cites.
Neural predictor for neural architecture search
Wei Wen, Hanxiao Liu, Yiran Chen, Hai Li, Gabriel Bender, and Pieter-Jan Kindermans · 2020
Later among the works it cites.
Neural architecture performance prediction using graph neural networks
Jovita Lukasik, David Friede, Heiner Stuckenschmidt, and Margret Keuper · 2020
Later among the works it cites.
What is being transferred in transfer learning?
Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang · 2020
Later among the works it cites.
Meta learning backpropagation and improving it
Louis Kirsch and Jürgen Schmidhuber · 2020
Later among the works it cites.
Towards fast adaptation of neural architectures with meta learning
Dongze Lian, Yin Zheng, Yintao Xu, Yanxiong Lu, Leyu Lin, Peilin Zhao, Junzhou Huang, and Shenghua Gao · 2020
Later among the works it cites.
Meta-learning of neural architectures for few-shot learning
Thomas Elsken, Benedikt Staffler, Jan Hendrik Metzen, and Frank Hutter · 2020
Later among the works it cites.
Catch: Context-based meta reinforcement learning for transferrable architecture search
Xin Chen, Yawen Duan, Zewei Chen, Hang Xu, Zihao Chen, Xiaodan Liang, Tong Zhang, and Zhenguo Li · 2020
Later among the works it cites.
Bignas: Scaling up neural architecture search with big single-stage models
Jiahui Yu, Pengchong Jin, Hanxiao Liu, Gabriel Bender, Pieter-Jan Kindermans, Mingxing Tan, Thomas Huang, Xiaodan Song, Ruoming Pang, and Quoc Le · 2020
Later among the works it cites.
A systematic survey on deep generative models for graph generation
Xiaojie Guo and Liang Zhao · 2020
Later among the works it cites.
Designing network design spaces
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár · 2020
Later among the works it cites.
Spagan: Shortest path graph attention network
Yiding Yang, Xinchao Wang, Mingli Song, Junsong Yuan, and Dacheng Tao · 2021
Closest in time.
Survey on graph embeddings and their applications to machine learning problems on graphs
Ilya Makarov, Dmitrii Kiselev, Nikita Nikitinsky, and Lovro Subelj · 2021
Closest in time.
Meta learning black-box population-based optimizers
Hugo Siqueira Gomes, Benjamin Léger, and Christian Gagné · 2021
Closest in time.
Automl: A survey of the state-of-the-art
Xin He, Kaiyong Zhao, and Xiaowen Chu · 2021
Closest in time.
Deep graph similarity learning: A survey
Guixiang Ma, Nesreen K Ahmed, Theodore L Willke, and S Yu Philip · 2021
Closest in time.
Gradinit: Learning to initialize neural networks for stable and efficient training
Chen Zhu, Renkun Ni, Zheng Xu, Kezhi Kong, W Ronny Huang, and Tom Goldstein · 2021
Closest in time.
Data-driven weight initialization with sylvester solvers
Debasmit Das, Yash Bhalgat, and Fatih Porikli · 2021
Closest in time.