Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have brought about significant transformations in human society.
Stochastic estimation of the maximum of a regression function
Jack Kiefer and Jacob Wolfowitz · 1952
Earlier work this paper cites.
A simplex method for function minimization
John A Nelder and Roger Mead · 1965
Earlier work this paper cites.
A stochastic approximation technique for generating maximum likelihood parameter estimates
James C Spall · 1987
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
James C Spall · 1992
Earlier work this paper cites.
Implementation of the simultaneous perturbation algorithm for stochastic optimization
James C Spall · 1998
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Optimal approximate matrix product in terms of stable rank
Michael B Cohen, Jelani Nelson, and David P Woodruff · 2015
Earlier work this paper cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
John C Duchi, Michael I Jordan, Martin J Wainwright, and Andre Wibisono · 2015
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh · 2017
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
signsgd via zeroth-order oracle
Sijia Liu, Pin-Yu Chen, Xiangyi Chen, and Mingyi Hong · 2018
Earlier work this paper cites.
Simple random search of static linear policies is competitive for reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Gradientless descent: High-dimensional zeroth-order optimization
Daniel Golovin, John Karro, Greg Kochanski, Chansoo Lee, Xingyou Song, and Qiuyi Zhang · 2019
Cited alongside, same era.
Camembert: a tasty french language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suarez, Yoann Dupont, Laurent Romary, Eric Villemonte de La Clergerie, Djame Seddah, and Benoit Sagot · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Algorithm and hardness for dynamic attention maintenance in large language models
Jan van den Brand, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Attention scheme inspired softmax regression
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
Randomized and deterministic attention sparsification algorithms for over-parameterized feature dimension
Yichuan Deng, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An admm based framework for automl pipeline configuration
Sijia Liu, Parikshit Ram, Deepak Vijaykeerthy, Djallel Bouneffouf, Gregory Bramble, Horst Samulowitz, Dakuo Wang, Andrew Conn, and Alexander Gray · 2020
Cited alongside, same era.
Attention-based sentiment analysis using convolutional and recurrent neural network
Mohd Usama, Belal Ahmad, Enmin Song, M Shamim Hossain, Mubarak Alrashoud, and Ghulam Muhammad · 2020
Cited alongside, same era.
Mongoose: A learnable lsh framework for efficient neural network training
Beidi Chen, Zichang Liu, Binghui Peng, Zhaozhuo Xu, Jonathan Lingjie Li, Tri Dao, Zhao Song, Anshumali Shrivastava, and Christopher Re · 2021
Cited alongside, same era.
Attention mechanism for neural machine translation: A survey
Weihua He, Yongyun Wu, and Xiaohua Li · 2021
Cited alongside, same era.
Approximating how single head attention learns
Charlie Snell, Ruiqi Zhong, Dan Klein, and Jacob Steinhardt · 2021
Cited alongside, same era.
Zeroth-order nonconvex stochastic optimization: Handling constraints, high dimensionality, and saddle points
Krishnakumar Balasubramanian and Saeed Ghadimi · 2022
Cited alongside, same era.
Optimizing language models for dialogue
ChatGPT · 2022
Cited alongside, same era.
Yeqi Gao, Zhao Song, and Shenghao Xie · 2023
Closest in time.
Differentially private attention computation
Yeqi Gao, Zhao Song, and Xin Yang · 2023
Closest in time.
An iterative algorithm for rescaled hyperbolic functions regression
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Closest in time.
How do transformers learn topic structure: Towards a mechanistic understanding
Yuchen Li, Yuanzhi Li, and Andrej Risteski · 2023
Closest in time.
The closeness of in-context learning and weight shifting for softmax regression
Shuai Li, Zhao Song, Yu Xia, Tong Yu, and Tianyi Zhou · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Fine-tuning language models with just forward passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Ritwik Sinha, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Infoprompt: Information-theoretic soft prompt tuning for natural language understanding
Junda Wu, Tong Yu, Rui Wang, Zhao Song, Ruiyi Zhang, Handong Zhao, Chaochao Lu, Shuai Li, and Ricardo Henao · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Closest in time.
Eric Zelikman, Qian Huang, Percy Liang, Nick Haber, and Noah D Goodman · 2023
Closest in time.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark W. Barrett, Zhangyang Wang, and Beidi Chen · 2023
Closest in time.