Fetching the paper…
Reading the bibliography…
Hypernetworks, neural networks that predict the parameters of another neural network, are powerful models that have been successfully used in diverse applications from image generation to multi-task learning.
Measures of the amount of ecologic association between species
Lee R Dice · 1945
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
A visual vocabulary for flower classification
M-E Nilsback and Andrew Zisserman · 2006
Earlier work this paper cites.
Open access series of imaging studies (oasis): cross-sectional mri data in young, middle aged, nondemented, and demented older adults
Daniel S Marcus, Tracy H Wang, Jamie Parker, John G Csernansky, John C Morris, and Randy L Buckner · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi, Benjamin Recht, et al · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas Le Roux, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V Le · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2016
Cited alongside, same era.
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
Sergey Ioffe · 2017
Cited alongside, same era.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
David Krueger, Chin-Wei Huang, Riashat Islam, Ryan Turner, Alexandre Lacoste, and Aaron Courville · 2017
Cited alongside, same era.
Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion
Joan Serrà, Santiago Pascual, and Carlos Segura · 2019
Later among the works it cites.
You only train once: Loss-conditional training of deep networks
Alexey Dosovitskiy and Josip Djolonga · 2020
Later among the works it cites.
Implicit neural representations with periodic activation functions
Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein · 2020
Later among the works it cites.
Fourier features let networks learn high frequency functions in low dimensional domains
Matthew Tancik, Pratul P Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T Barron, and Ren Ng · 2020
Later among the works it cites.
Continual learning with hypernetworks
Johannes von Oswald, Christian Henning, Benjamin F. Grewe, and João Sacramento · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Implicit weight uncertainty in neural networks
Nick Pawlowski, Andrew Brock, Matthew CH Lee, Martin Rajchl, and Ben Glocker · 2017
Cited alongside, same era.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
Understanding batch normalization
Nils Bjorck, Carla P Gomes, Bart Selman, and Kilian Q Weinberger · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Cited alongside, same era.
Meta-learning via hypernetworks
Dominic Zhao, Johannes von Oswald, Seijin Kobayashi, João Sacramento, and Benjamin F Grewe · 2020
Later among the works it cites.
Meta-learning symmetries by reparameterization
Allan Zhou, Tom Knowles, and Chelsea Finn · 2020
Later among the works it cites.
Meta internal learning
Raphael Bensadoun, Shir Gur, Tomer Galanti, and Lior Wolf · 2021
Later among the works it cites.
Continual learning in recurrent neural networks
Benjamin Ehret, Christian Henning, Maria R. Cervera, Alexander Meulemans, Johannes von Oswald, and Benjamin F. Grewe · 2021
Later among the works it cites.
Posterior meta-replay for continual learning
Christian Henning, Maria R. Cervera, Francesco D’Angelo, Johannes von Oswald, Regina Traber, Benjamin Ehret, Seijin Kobayashi, Benjamin F. Grewe, and João Sacramento · 2021
Later among the works it cites.
Beyond batchnorm: Towards a unified understanding of normalization in deep learning
Ekdeep S Lubana, Robert Dick, and Hidenori Tanaka · 2021
Later among the works it cites.
Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2021
Later among the works it cites.
Hypergrid transformers: Towards a single model for multiple tasks
Yi Tay, Zhe Zhao, Dara Bahri, Don Metzler, and Da-Cheng Juan · 2021
Later among the works it cites.
Regularization-agnostic compressed sensing mri reconstruction with hypernetworks
Alan Q Wang, Adrian V Dalca, and Mert R Sabuncu · 2021
Later among the works it cites.
Hyperstyle: Stylegan inversion with hypernetworks for real image editing
Yuval Alaluf, Omer Tov, Ron Mokady, Rinon Gal, and Amit Bermano · 2022
Later among the works it cites.
Multi-rate vae: Train once, get the full rate-distortion curve
Juhan Bae, Michael R Zhang, Michael Ruan, Eric Wang, So Hasegawa, Jimmy Ba, and Roger Grosse · 2022
Later among the works it cites.
Hyperinverter: Improving stylegan inversion via hypernetwork
Tan M Dinh, Anh Tuan Tran, Rang Nguyen, and Binh-Son Hua · 2022
Later among the works it cites.
Learning the effect of registration hyperparameters with hypermorph
Andrew Hoopes, Malte Hoffman, Douglas N. Greve, Bruce Fischl, John Guttag, and Adrian V. Dalca · 2022
Later among the works it cites.
Amortized learning of dynamic feature scaling for image segmentation
Jose Javier Gonzalez Ortiz, John Guttag, and Adrian V. Dalca · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models, 2023
Lvmin Zhang and Maneesh Agrawala · 2023
Closest in time.