Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Automated Machine Learning: Methods, Systems, Challenges
Frank Hutter, Lars Kotthoff, and Joaquin Vanschoren, editors. 2018 · 2018
Later among the works it cites.
An alternative view: When does SGD escape local minima?
Bobby Kleinberg, Yuanzhi Li, and Yang Yuan. 2018 · 2018
Later among the works it cites.
Regularized evolution for image classifier architecture search
Original
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V. Le. 2018 · 2018
Later among the works it cites.
On the convergence of adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar. 2018 · 2018
Later among the works it cites.
Validating quantum computers using randomized model circuits
Andrew W. Cross, Lev S. Bishop, Sarah Sheldon, Paul D. Nation, and Jay M. Gambetta. 2019 · 2019
Later among the works it cites.
Neural architecture search: A survey
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. 2019 · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2019 · 2019
Later among the works it cites.
On the bounds of function approximations
Adrian de Wynter. 2019 · 2019
Later among the works it cites.
Language models are few-shot learners
Original
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Closest in time.
Geometry-aware gradient algorithms for neural architecture search
Original
Liam Li, Mikhail Khodak, Maria-Florina Balcan, and Ameet Talwalkar. 2020 · 2020
Closest in time.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. 2020 · 2020
Closest in time.
A unified convergence analysis for shuffling-type gradient methods
Original
Lam M. Nguyen, Quoc Tran-Dinh, Dzung T. Phan, Phuong Ha Nguyen, and Marten van Dijk. 2020 · 2020
Closest in time.
Comparing rewinding and fine-tuning in neural network pruning
Alex Renda, Jonathan Frankle, and Michael Carbin. 2020 · 2020
Closest in time.
Turing-NLG: A 17-billion-parameter language model by Microsoft
Corby Rosset. 2020 · 2020
Closest in time.
On complexity of finding stationary points of nonsmooth nonconvex functions
Original
Jingzhao Zhang, Hongzhou Lin, Stefanie Jegelka, Ali Jadbabaie, and Suvrit Sra. 2020 · 2020
Closest in time.