Fetching the paper…
Reading the bibliography…
In the realm of deep learning, transformers have emerged as a dominant architecture, particularly in natural language processing tasks.
Inversion of feedforward neural networks: algorithms and applications
Craig A Jensen, Russell D Reed, Robert Jackson Marks, Mohamed A El-Sharkawi, Jae-Byung Jung, Robert T Miyamoto, Gregory M Anderson, and Christian J Eggen · 1999
Earlier work this paper cites.
Inverting feedforward neural networks using linear and nonlinear programming
Bao-Liang Lu, Hajime Kita, and Yoshikazu Nishikawa · 1999
Earlier work this paper cites.
The volumetric barrier for semidefinite programming
Kurt M Anstreicher · 2000
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Understanding deep image representations by inverting them
Aravindh Mahendran and Andrea Vedaldi · 2015
Earlier work this paper cites.
Inverting visual representations with convolutional networks
Alexey Dosovitskiy and Thomas Brox · 2016
Earlier work this paper cites.
Multi-way, multilingual neural machine translation with a shared attention mechanism
Orhan Firat, Kyunghyun Cho, and Yoshua Bengio · 2016
Earlier work this paper cites.
Defeating image obfuscation with deep learning
Richard McPherson, Reza Shokri, and Vitaly Shmatikov · 2016
Earlier work this paper cites.
The limitations of deep learning in adversarial settings
Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami · 2016
Earlier work this paper cites.
Deep models under the gan: information leakage from collaborative deep learning
Briland Hitaj, Giuseppe Ateniese, and Fernando Perez-Cruz · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Fine-grained attention mechanism for neural machine translation
Heeyoul Choi, Kyunghyun Cho, and Yoshua Bengio · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Earlier work this paper cites.
Gram-gauss-newton method: Learning overparameterized neural networks for regression problems
Tianle Cai, Ruiqi Gao, Jikai Hou, Siyu Chen, Dong Wang, Di He, Zhihua Zhang, and Liwei Wang · 2019
Earlier work this paper cites.
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Earlier work this paper cites.
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Earlier work this paper cites.
Camembert: a tasty french language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suarez, Yoann Dupont, Laurent Romary, Eric Villemonte de La Clergerie, Djame Seddah, and Benoit Sagot · 2019
Earlier work this paper cites.
Exploiting unintended feature leakage in collaborative learning
Luca Melis, Congzheng Song, Emiliano De Cristofaro, and Vitaly Shmatikov · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Quadratic suffices for over-parametrization via matrix chernoff bound
Zhao Song and Xin Yang · 2019
Earlier work this paper cites.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Earlier work this paper cites.
Deep leakage from gradients
Ligeng Zhu, Zhijian Liu, and Song Han · 2019
Earlier work this paper cites.
Fast convergence of natural gradient descent for over-parameterized neural networks
Guodong Zhang, James Martens, and Roger B Grosse · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Training (overparametrized) neural networks in near-linear time
Jan van den Brand, Binghui Peng, Zhao Song, and Omri Weinstein · 2020
Earlier work this paper cites.
Instahide: Instance-hiding schemes for private distributed learning
Yangsibo Huang, Zhao Song, Kai Li, and Sanjeev Arora · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Earlier work this paper cites.
Generalized leverage score sampling for neural networks
Jason D Lee, Ruoqi Shen, Zhao Song, Mengdi Wang, et al · 2020
Earlier work this paper cites.
Transformer based deep intelligent contextual embedding for twitter sentiment analysis
Usman Naseem, Imran Razzak, Katarzyna Musial, and Muhammad Imran · 2020
Earlier work this paper cites.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Cited alongside, same era.
Privacy risks of general-purpose language models
Xudong Pan, Mi Zhang, Shouling Ji, and Min Yang · 2020
Cited alongside, same era.
A survey of privacy attacks in machine learning
Maria Rigaki and Sebastian Garcia · 2020
Cited alongside, same era.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2020
Cited alongside, same era.
Attention-based sentiment analysis using convolutional and recurrent neural network
Mohd Usama, Belal Ahmad, Enmin Song, M Shamim Hossain, Mubarak Alrashoud, and Ghulam Muhammad · 2020
Cited alongside, same era.
How to protect copyright data in optimization of large language models?
Timothy Chu, Zhao Song, and Chiwun Yang · 2023
Closest in time.
Zero-th order algorithm for softmax attention optimization
Yichuan Deng, Zhihang Li, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Attention scheme inspired softmax regression
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
Randomized and deterministic attention sparsification algorithms for over-parameterized feature dimension
Yichuan Deng, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Convergence of two-layer regression with nonlinear units
Yichuan Deng, Zhao Song, and Shenghao Xie · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenqi Wei, Ling Liu, Margaret Loper, Ka-Ho Chow, Mehmet Emre Gursoy, Stacey Truex, and Yanzhao Wu · 2020
Cited alongside, same era.
The secret revealer: Generative model-inversion attacks against deep neural networks
Yuheng Zhang, Ruoxi Jia, Hengzhi Pei, Wenxiao Wang, Bo Li, and Dawn Song · 2020
Cited alongside, same era.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Cited alongside, same era.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Cited alongside, same era.
idlg: Improved deep leakage from gradients
Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen · 2020
Cited alongside, same era.
Over-parameterized adversarial training: An analysis overcoming the curse of dimensionality
Yi Zhang, Orestis Plevrakis, Simon S Du, Xingguo Li, Zhao Song, and Sanjeev Arora · 2020
Cited alongside, same era.
A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization
HanQin Cai, Yuchen Lou, Daniel Mckenzie, and Wotao Yin · 2021
Cited alongside, same era.
Closest in time.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Yeqi Gao, Zhao Song, and Shenghao Xie · 2023
Closest in time.
Yeqi Gao, Zhao Song, and Shenghao Xie · 2023
Closest in time.
Differentially private attention computation
Yeqi Gao, Zhao Song, and Xin Yang · 2023
Closest in time.
Gradientcoin: A peer-to-peer decentralized large language models
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Closest in time.
Fast quantum algorithm for attention computation
Yeqi Gao, Zhao Song, Xin Yang, and Ruizhe Zhang · 2023
Closest in time.
A nearly-linear time algorithm for structured support vector machines
Yuzhou Gu, Zhao Song, and Lichen Zhang · 2023
Closest in time.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Closest in time.
Polysketchformer: Fast transformers via sketches for polynomial kernels
Praneeth Kacham, Vahab Mirrokni, and Peilin Zhong · 2023
Closest in time.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma · 2023
Closest in time.
How do transformers learn topic structure: Towards a mechanistic understanding
Yuchen Li, Yuanzhi Li, and Andrej Risteski · 2023
Closest in time.
The closeness of in-context learning and weight shifting for softmax regression
Shuai Li, Zhao Song, Yu Xia, Tong Yu, and Tianyi Zhou · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Fine-tuning language models with just forward passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora · 2023
Closest in time.
A kernel-based view of language model fine-tuning
Sadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen, and Sanjeev Arora · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Trainable transformer in transformer
Abhishek Panigrahi, Sadhika Malladi, Mengzhou Xia, and Sanjeev Arora · 2023
Closest in time.
Task-specific skill localization in fine-tuned language models
Abhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, and Sanjeev Arora · 2023
Closest in time.
Efficient sgd neural network training via sublinear activated neuron identification
Lianke Qin, Zhao Song, and Yuanyuan Yang · 2023
Closest in time.
Chatgpt: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope
Partha Pratim Ray · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D.Manning, and Chelsea Finn · 2023
Closest in time.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2023
Closest in time.
Do pretrained transformers really learn in-context by gradient descent?
Lingfeng Shen, Aayush Mishra, and Daniel Khashabi · 2023
Closest in time.
Ritwik Sinha, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Differentially private non-convex learning for multi-layer neural networks
Hanpu Shen, Cheng-Long Wang, Zihang Xiang, Yiming Ying, and Di Wang · 2023
Closest in time.
Solving attention kernel regression problem via pre-conditioner
Zhao Song, Junze Yin, and Lichen Zhang · 2023
Closest in time.
Transformers as support vector machines
Davoud Ataee Tarzanagh, Yingcong Li, Christos Thrampoulidis, and Samet Oymak · 2023
Closest in time.
Provable copyright protection for generative models
Nikhil Vyas, Sham Kakade, and Boaz Barak · 2023
Closest in time.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Closest in time.
Reconstructing training data from model gradient, provably
Zihan Wang, Jason Lee, and Qi Lei · 2023
Closest in time.
Sheared llama: Accelerating language model pre-training via structured pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen · 2023
Closest in time.
Federated learning of gboard language models with differential privacy
Zheng Xu, Yanxiang Zhang, Galen Andrew, Christopher A Choquette-Choo, Peter Kairouz, H Brendan McMahan, Jesse Rosenstock, and Yuanbo Zhang · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Closest in time.