Fetching the paper…
Reading the bibliography…
Bi-level optimization (BO) has become a fundamental mathematical framework for addressing hierarchical machine learning problems.
A simple automatic derivative evaluation program
R. E. Wengert · 1964
Earlier work this paper cites.
Numerical analysis of spectral methods: theory and applications
David Gottlieb and Steven A Orszag · 1977
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J. Williams and David Zipser · 1989
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A. Pearlmutter · 1994
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
The implicit function theorem: History, theory, and applications
Steven George Krantz and Harold R Parks · 2002
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
Spectral methods: fundamentals in single domains
Claudio Canuto, M Yousuff Hussaini, Alfio Quarteroni, and Thomas A Zang · 2007
Earlier work this paper cites.
An overview of bilevel optimization
Benoît Colson, Patrice Marcotte, and Gilles Savard · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Conjugate gradient method
John L Nazareth · 2009
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio · 2013
Earlier work this paper cites.
Numerical methods for partial differential equations
William F Ames · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
John C Duchi, Michael I Jordan, Martin J Wainwright, and Andre Wibisono · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Earlier work this paper cites.
Scalable gradient-based tuning of continuous regularization hyperparameters
Jelena Luketina, Mathias Berglund, Klaus Greff, and Tapani Raiko · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Forward and reverse gradient-based hyperparameter optimization
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2017
Earlier work this paper cites.
A review on bilevel optimization: From classical to evolutionary approaches and applications
Ankur Sinha, Pekka Malo, and Kalyanmoy Deb · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina N. Toutanova · 2018
Earlier work this paper cites.
Bilevel programming for hyperparameter optimization and meta-learning
Luca Franceschi, Paolo Frasconi, Saverio Salzo, Riccardo Grazzi, and Massimiliano Pontil · 2018
Cited alongside, same era.
Approximation methods for bilevel programming
Saeed Ghadimi and Mengdi Wang · 2018
Cited alongside, same era.
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2018
Cited alongside, same era.
Zeroth-order stochastic variance reduction for nonconvex optimization
Sijia Liu, Bhavya Kailkhura, Pin-Yu Chen, Paishun Ting, Shiyu Chang, and Lisa Amini · 2018
Cited alongside, same era.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Cited alongside, same era.
Physics-informed neural networks with hard constraints for inverse design
Lu Lu, Raphael Pestourie, Wenjie Yao, Zhicheng Wang, Francesc Verdugo, and Steven G Johnson · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
Learning by directional gradient descent
David Silver, Anirudh Goyal, Ivo Danihelka, Matteo Hessel, and Hado van Hasselt · 2021
Later among the works it cites.
Robust watermarking for deep neural networks via bi-level optimization
Peng Yang, Yingjie Lao, and Ping Li · 2021
Later among the works it cites.
Adversarial regularization as stackelberg game: An unrolled optimization approach
Simiao Zuo, Chen Liang, Haoming Jiang, Xiaodong Liu, Pengcheng He, Jianfeng Gao, Weizhu Chen, and Tuo Zhao · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Numerical methods for solving partial differential equations: a comprehensive introduction for scientists and engineers
George F Pinder · 2018
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V. Ugur Güney, Yann N. Dauphin, and Léon Bottou · 2018
Cited alongside, same era.
Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros · 2018
Cited alongside, same era.
Neural architecture search: A survey
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter · 2019
Cited alongside, same era.
Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations
Maziar Raissi, Paris Perdikaris, and George E Karniadakis · 2019
Cited alongside, same era.
Meta-learning with implicit gradients
Aravind Rajeswaran, Chelsea Finn, Sham M Kakade, and Sergey Levine · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Atılım Güneş Baydin, Barak A Pearlmutter, Don Syme, Frank Wood, and Philip Torr · 2022
Later among the works it cites.
Privacy for free: How does dataset condensation help privacy?
Tian Dong, Bo Zhao, and Lingjuan Lyu · 2022
Later among the works it cites.
Zhongkai Hao, Chengyang Ying, Hang Su, Jun Zhu, Jian Song, and Ze Cheng · 2022
Later among the works it cites.
Streamingqa: A benchmark for adaptation to new knowledge over time in question answering models
Adam Liska, Tomas Kocisky, Elena Gribovskaya, Tayfun Terzi, Eren Sezener, Devang Agrawal, D’Autume Cyprien De Masson, Tim Scholtes, Manzil Zaheer, Susannah Young, et al · 2022
Later among the works it cites.
Bome! bilevel optimization made easy: A simple first-order approach
Bo Liu, Mao Ye, Stephen Wright, Peter Stone, and Qiang Liu · 2022
Later among the works it cites.
Scaling forward gradient with local losses
Mengye Ren, Simon Kornblith, Renjie Liao, and Geoffrey Hinton · 2022
Later among the works it cites.
Cafe: Learning to condense dataset by aligning features
Kai Wang, Bo Zhao, Xiangyu Peng, Zheng Zhu, Shuo Yang, Shuo Wang, Guan Huang, Hakan Bilen, Xinchao Wang, and Yang You · 2022
Later among the works it cites.
Advancing model pruning via bi-level optimization
Yihua Zhang, Yuguang Yao, Parikshit Ram, Pu Zhao, Tianlong Chen, Mingyi Hong, Yanzhi Wang, and Sijia Liu · 2022
Later among the works it cites.
Revisiting and advancing fast adversarial training through the lens of bi-level optimization
Yihua Zhang, Guanhua Zhang, Prashant Khanduri, Mingyi Hong, Shiyu Chang, and Sijia Liu · 2022
Later among the works it cites.
Near-optimal fully first-order algorithms for finding stationary points in bilevel optimization
Lesi Chen, Yaohua Ma, and Jingzhao Zhang · 2023
Later among the works it cites.
Nyström method for accurate and scalable implicit differentiation
Ryuichiro Hataya and Makoto Yamada · 2023
Later among the works it cites.
Meta-learning online adaptation of language models
Nathan Hu, Eric Mitchell, Christopher D Manning, and Chelsea Finn · 2023
Later among the works it cites.
A fully first-order method for stochastic bilevel optimization
Jeongyeol Kwon, Dohyun Kwon, Stephen Wright, and Robert D Nowak · 2023
Later among the works it cites.
Fine-tuning language models with just forward passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora · 2023
Later among the works it cites.
On penalty-based bilevel gradient descent method
Han Shen and Tianyi Chen · 2023
Later among the works it cites.
Dataset distillation: A comprehensive review
Ruonan Yu, Songhua Liu, and Xinchao Wang · 2023
Later among the works it cites.
Yihua Zhang, Prashant Khanduri, Ioannis Tsaknakis, Yuguang Yao, Mingyi Hong, and Sijia Liu · 2023
Later among the works it cites.
Making scalable meta learning practical
Sang Choe, Sanket Vaibhav Mehta, Hwijeen Ahn, Willie Neiswanger, Pengtao Xie, Emma Strubell, and Eric Xing · 2024
Closest in time.
Picprop: Physics-informed confidence propagation for uncertainty quantification
Qianli Shen, Wai Hoh Tang, Zhun Deng, Apostolos Psaros, and Kenji Kawaguchi · 2024
Closest in time.
Online adaptation of language models with a memory of amortized contexts
Jihoon Tack, Jaehyung Kim, Eric Mitchell, Jinwoo Shin, Yee Whye Teh, and Jonathan Richard Schwarz · 2024
Closest in time.
Revisiting zeroth-order optimization for memory-efficient llm fine-tuning: A benchmark
Yihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li, Yimeng Zhang, Wenqing Zheng, Pin-Yu Chen, Jason D Lee, Wotao Yin, Mingyi Hong, et al · 2024
Closest in time.