High-dimensional statistics: A non-asymptotic viewpoint
Martin J Wainwright · 2019
Later among the works it cites.
Global convergence of adaptive gradient methods for an over-parameterized neural network
Original
Xiaoxia Wu, Simon S Du, and Rachel Ward · 2019
Later among the works it cites.
Fast convergence of natural gradient descent for over-parameterized neural networks
Guodong Zhang, James Martens, and Roger B Grosse · 2019
Later among the works it cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
An improved cutting plane method for convex optimization, convex-concave games and its applications
Haotian Jiang, Yin Tat Lee, Zhao Song, and Sam Chiu-wai Wong · 2020
Later among the works it cites.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Training (overparametrized) neural networks in near-linear time
Original
Jan van den Brand, Binghui Peng, Zhao Song, and Omri Weinstein · 2021
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Later among the works it cites.
Oblivious sketching-based central path method for linear programming
Zhao Song and Zheng Yu · 2021
Later among the works it cites.
Does preprocessing help training over-parameterized neural networks?
Zhao Song, Shuo Yang, and Ruizhe Zhang · 2021
Later among the works it cites.
Optimizing language models for dialogue
ChatGPT · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Original
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
Solving sdp faster: A robust ipm framework and efficient implementation
Baihe Huang, Shunhua Jiang, Zhao Song, Runzhou Tao, and Ruizhe Zhang · 2022
Closest in time.
Sustainable ai: Environmental implications, challenges and opportunities
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al · 2022
Closest in time.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Closest in time.
Opt: Open pre-trained transformer language models
Original
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Closest in time.
Convergence of two-layer regression with nonlinear units
Original
Yichuan Deng, Zhao Song, and Shenghao Xie · 2023
Closest in time.
Gpt-4 technical report
Original
OpenAI · 2023
Closest in time.
Enhancing stochastic gradient descent: A unified framework and novel acceleration methods for faster convergence
Original
Yichuan Deng, Zhao Song, and Chiwun Yang · 2024
Closest in time.