Fetching the paper…
Reading the bibliography…
Diffusion models have led to significant advancements in generative modelling.
Generative modeling by estimating gradients of the data distribution, 2020
Yang Song and Stefano Ermon · 1907
Earlier work this paper cites.
Extensions of lipschitz maps into banach spaces
William B Johnson, Joram Lindenstrauss, and Gideon Schechtman · 1986
Earlier work this paper cites.
A theoretical framework for back-propagation
Yann LeCun, D Touresky, G Hinton, and T Sejnowski · 1988
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
On “natural” learning and pruning in multilayered perceptrons
Tom Heskes · 2000
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
An elementary proof of a theorem of Johnson and Lindenstrauss
Sanjoy Dasgupta and Anupam Gupta · 2003
Earlier work this paper cites.
The Implicit Function Theorem
Steven G. Krantz and Harold R. Parks · 2003
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Deep Unsupervised Learning using Nonequilibrium Thermodynamics, November 2015
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
A Kronecker-factored approximate Fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Earlier work this paper cites.
Practical Gauss-Newton optimisation for deep learning
Aleksandar Botev, Hippolyt Ritter, and David Barber · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
Exact natural gradient in deep linear networks and its application to the nonlinear case
Alberto Bernacchia, Mate Lengyel, and Guillaume Hennequin · 2018
Earlier work this paper cites.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Earlier work this paper cites.
On the Accuracy of Influence Functions for Measuring Group Effects, November 2019
Pang Wei Koh, Kai-Siang Ang, Hubert H. K. Teo, and Percy Liang · 2019
Cited alongside, same era.
Limitations of the empirical Fisher approximation for natural gradient descent
Frederik Kunstner, Lukas Balles, and Philipp Hennig · 2019
Cited alongside, same era.
Which algorithmic choices matter at which batch sizes? Insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George E. Dahl, Christopher J. Shallue, and Roger B. Grosse · 2019
Cited alongside, same era.
Relatif: Identifying explanatory training samples via relative influence
Elnaz Barshan, Marc-Etienne Brunet, and Gintare Karolina Dziugaite · 2020
Cited alongside, same era.
On second-order group influence functions for black-box predictions
Samyadeep Basu, Xuchen You, and Soheil Feizi · 2020
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models, 2022
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev · 2022
Later among the works it cites.
Denoising Diffusion Implicit Models, October 2022
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2022
Later among the works it cites.
Kronecker-Factored Approximate Curvature for modern neural network architectures
Runa Eschenhagen, Alexander Immer, Richard E. Turner, Frank Schneider, and Philipp Hennig · 2023
Later among the works it cites.
The Journey, Not the Destination: How Data Guides Diffusion Models, December 2023
Kristian Georgiev, Joshua Vendrow, Hadi Salman, Sung Min Park, and Aleksander Madry · 2023
Later among the works it cites.
Studying Large Language Model Generalization with Influence Functions, August 2023
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamilė Lukošiūtė, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Samuel R. Bowman · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2020
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
KAISA: an adaptive second-order optimizer framework for deep neural networks
J. Gregory Pauloski, Qi Huang, Lei Huang, Shivaram Venkataraman, Kyle Chard, Ian T. Foster, and Zhao Zhang · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
SKFAC: Training neural networks with faster Kronecker-factored approximate curvature
Zedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li, Yue Wu, Fan Yu, Zidong Wang, and Min Wang · 2021
Cited alongside, same era.
If Influence Functions are the Answer, Then What is the Question?, September 2022
Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi, and Roger Grosse · 2022
Cited alongside, same era.
Later among the works it cites.
Variational Diffusion Models, April 2023
Diederik P. Kingma, Tim Salimans, Ben Poole, and Jonathan Ho · 2023
Later among the works it cites.
DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion Models
Yongchan Kwon, Eric Wu, Kevin Wu, and James Zou · 2023
Later among the works it cites.
TRAK: Attributing Model Behavior at Scale, April 2023
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry · 2023
Later among the works it cites.
Training data attribution via approximate unrolled differentiation, 2024
Juhan Bae, Wu Lin, Jonathan Lorraine, and Roger Grosse · 2024
Closest in time.
Generalization in diffusion models arises from geometry-adaptive harmonic representations, April 2024
Zahra Kadkhodaie, Florentin Guth, Eero P. Simoncelli, and Stéphane Mallat · 2024
Closest in time.
Diffusion attribution score: Evaluating training data influence in diffusion model
Jinxu Lin, Linwei Tao, Minjing Dong, and Chang Xu · 2024
Closest in time.
Image generator litigation
Joseph Saveri and Matthew Butterick · 2024
Closest in time.
Language model litigation
Joseph Saveri and Matthew Butterick · 2024
Closest in time.
Denoising diffusion probabilistic models in six simple steps, 2024
Richard E. Turner, Cristiana-Diana Diaconu, Stratis Markou, Aliaksandra Shysheya, Andrew Y. K. Foong, and Bruno Mlodozeniec · 2024
Closest in time.
Intriguing Properties of Data Attribution on Diffusion Models, March 2024
Xiaosen Zheng, Tianyu Pang, Chao Du, Jing Jiang, and Min Lin · 2024
Closest in time.
Position: Curvature matrices should be democratized via linear operators
Felix Dangel, Runa Eschenhagen, Weronika Ormaniec, Andres Fernandez, Lukas Tatzel, and Agustinus Kristiadi · 2025
Closest in time.