Fetching the paper…
Reading the bibliography…
Alternatives to backpropagation have long been studied to better understand how biological brains may learn.
Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences
P. J. Werbos · 1974
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Competitive learning: From interactive activation to adaptive resonance
Stephen Grossberg · 1987
Earlier work this paper cites.
The recent excitement about neural networks
Francis Crick · 1989
Earlier work this paper cites.
Contrastive hebbian learning in the continuous hopfield model
Javier R Movellan · 1991
Earlier work this paper cites.
Biologically plausible error-driven learning using local activation differences: The generalized recirculation algorithm
Randall C O’Reilly · 1996
Earlier work this paper cites.
Random synaptic feedback weights support error backpropagation for deep learning
Timothy P Lillicrap, Daniel Cownden, Douglas B Tweed, and Colin J Akerman · 2016
Earlier work this paper cites.
Direct feedback alignment provides learning in deep neural networks
Arild Nøkland · 2016
Earlier work this paper cites.
Decoupled neural interfaces using synthetic gradients
Max Jaderberg, Wojciech Marian Czarnecki, Simon Osindero, Oriol Vinyals, Alex Graves, David Silver, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Assessing the scalability of biologically-motivated deep learning algorithms and architectures
Sergey Bartunov, Adam Santoro, Blake Richards, Luke Marris, Geoffrey E Hinton, and Timothy Lillicrap · 2018
Earlier work this paper cites.
A 0.086-mm 2 12.7-pj/sop 64k-synapse 256-neuron online-learning digital spiking neuromorphic processor in 28-nm cmos
Charlotte Frenkel, Martin Lefebvre, Jean-Didier Legat, and David Bol · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Training neural networks with local error signals
Arild Nøkland and Lars Hiller Eidnes · 2019
Earlier work this paper cites.
Using memristors for robust local learning of hardware restricted boltzmann machines
Maxence Ernoult, Julie Grollier, and Damien Querlioz · 2019
Earlier work this paper cites.
Greedy layerwise learning can scale to imagenet
Eugene Belilovsky, Michael Eickenberg, and Edouard Oyallon · 2019
Earlier work this paper cites.
Principled training of neural networks with direct feedback alignment
Julien Launay, Iacopo Poli, and Florent Krzakala · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Backpropagation and the brain
Timothy P Lillicrap, Adam Santoro, Luke Marris, Colin J Akerman, and Geoffrey Hinton · 2020
Cited alongside, same era.
Direct feedback alignment scales to modern deep learning tasks and architectures
Julien Launay, Iacopo Poli, François Boniface, and Florent Krzakala · 2020
Cited alongside, same era.
Hardware beyond backpropagation: a photonic co-processor for direct feedback alignment
Julien Launay, Iacopo Poli, Kilian Müller, Gustave Pariente, Igor Carron, Laurent Daudet, Florent Krzakala, and Sylvain Gigan · 2020
Cited alongside, same era.
Bottom-up and top-down neuromorphic processor design: Unveiling roads to embedded cognition
Charlotte Frenkel · 2020
Cited alongside, same era.
Light-in-the-loop: using a photonics co-processor for scalable training of neural networks
Julien Launay, Iacopo Poli, Kilian Müller, Igor Carron, Laurent Daudet, Florent Krzakala, and Sylvain Gigan · 2020
Cited alongside, same era.
Photonic differential privacy with direct feedback alignment
Ruben Ohana, Hamlet Medina, Julien Launay, Alessandro Cappelli, Iacopo Poli, Liva Ralaivola, and Alain Rakotomamonjy · 2021
Later among the works it cites.
Do transformer modifications transfer across implementations and applications?
Sharan Narang, Hyung Won Chung, Yi Tay, Liam Fedus, Thibault Févry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, et al · 2021
Later among the works it cites.
Is the number of trainable parameters all that actually matters?
Amélie Chatelain, Amine Djeghri, Daniel Hesslow, and Julien Launay · 2022
Closest in time.
Data scaling laws in nmt: The effect of noise and architecture
Yamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang, Colin Cherry, Behnam Neyshabur, and Orhan Firat · 2022
Closest in time.
What language model architecture and pretraining objective work best for zero-shot generalization?
Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Laskin, Luke Metz, Seth Nabarro, Mark Saroufim, Badreddine Noune, Carlo Luschi, Jascha Sohl-Dickstein, and Pieter Abbeel · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Cited alongside, same era.
Ccnet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Édouard Grave · 2020
Cited alongside, same era.
Differentially private deep learning with direct feedback alignment
Jaewoo Lee and Daniel Kifer · 2020
Cited alongside, same era.
Learning without feedback: Fixed random learning signals allow for feedforward training of deep neural networks
Charlotte Frenkel, Martin Lefebvre, and David Bol · 2021
Cited alongside, same era.
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
Revisiting neural scaling laws in language and vision
Ibrahim Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai · 2022
Closest in time.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Closest in time.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S Morcos · 2022
Closest in time.
Reducing activation recomputation in large transformer models
Vijay Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro · 2022
Closest in time.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Closest in time.
Neuromorphic photonic technologies and architectures: scaling opportunities and performance frontiers
George Dabos, Dimitris V Bellas, Ripalta Stabile, Miltiadis Moralis-Pegios, George Giamougiannis, Apostolos Tsakyridis, Angelina Totovic, Elefterios Lidorikis, and Nikos Pleros · 2022
Closest in time.
Adversarial robustness by design through analog computing and synthetic gradients
Alessandro Cappelli, Ruben Ohana, Julien Launay, Laurent Meunier, Iacopo Poli, and Florent Krzakala · 2022
Closest in time.
Self-supervised pretraining for differentially private learning
Arash Asadian, Evan Weidner, and Lei Jiang · 2022
Closest in time.
What language model to train if you have one million gpu hours?
Teven Le Scao, Thomas Wang, Daniel Hesslow, Lucile Saulnier, Stas Bekman, M Saiful Bari, Stella Biderman, Hady Elsahar, Jason Phang, Ofir Press, et al · 2022
Closest in time.
Unified scaling laws for routed language models
Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, et al · 2022
Closest in time.
Scaling laws vs model architectures: How does inductive bias influence scaling?
Yi Tay, Mostafa Dehghani, Samira Abnar, Hyung Won Chung, William Fedus, Jinfeng Rao, Sharan Narang, Vinh Q Tran, Dani Yogatama, and Donald Metzler · 2022
Closest in time.