Fetching the paper…
Reading the bibliography…
Transformers have emerged as the leading architecture in deep learning, proving to be versatile and highly effective across diverse domains beyond language and image processing.
Improved Precision and Recall Metric for Assessing Generative Models
Kynkäänniemi, T.; Karras, T.; Laine, S.; Lehtinen, J.; and Aila, T. 2019 · 1904
Earlier work this paper cites.
Gram-gauss-newton method: Learning overparameterized neural networks for regression problems
Cai, T.; Gao, R.; Hou, J.; Chen, S.; Wang, D.; He, D.; Zhang, Z.; and Wang, L. 2019 · 1905
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019a · 1905
Earlier work this paper cites.
BERT Rediscovers the Classical NLP Pipeline
Tenney, I.; Das, D.; and Pavlick, E. 2019 · 1905
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 1905
Earlier work this paper cites.
What Does BERT Look At? An Analysis of BERT’s Attention
Clark, K.; Khandelwal, U.; Levy, O.; and Manning, C. D. 2019b · 1906
Earlier work this paper cites.
Quantum Entropy Scoring for Fast Robust Mean Estimation and Improved Outlier Detection
Dong, Y.; Hopkins, S. B.; and Li, J. 2019 · 1906
Earlier work this paper cites.
Quadratic suffices for over-parametrization via matrix chernoff bound
Song, Z.; and Yang, X. 2019 · 1906
Earlier work this paper cites.
Analyzing the Structure of Attention in a Transformer Language Model
Vig, J.; and Belinkov, Y. 2019 · 1906
Earlier work this paper cites.
Designing and Interpreting Probes with Control Tasks
Hewitt, J.; and Liang, P. 2019 · 1909
Earlier work this paper cites.
Ji, Z.; and Telgarsky, M. 2019 · 1909
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O (1/k2)
Nesterov, Y. 1983 · 1983
Earlier work this paper cites.
Optimal Brain Damage
LeCun, Y.; Denker, J.; and Solla, S. 1989 · 1989
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vapnik, V. 1991 · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T.; and Juditsky, A. B. 1992 · 1992
Earlier work this paper cites.
Optimal Brain Surgeon and general network pruning
Hassibi, B.; Stork, D.; and Wolff, G. 1993 · 1993
Earlier work this paper cites.
Building a Large Annotated Corpus of English: The Penn Treebank
Marcus, M. P.; Santorini, B.; and Marcinkiewicz, M. A. 1993 · 1993
Earlier work this paper cites.
The Volumetric Barrier for Semidefinite Programming
Anstreicher, K. M. 2000 · 2000
Earlier work this paper cites.
Training v-support vector classifiers: theory and algorithms
Chang, C.-C.; and Lin, C.-J. 2001 · 2001
Earlier work this paper cites.
An Improved Cutting Plane Method for Convex Optimization, Convex-Concave Games and its Applications
Jiang, H.; Lee, Y. T.; Song, Z.; and wai Wong, S. C. 2020b · 2004
Earlier work this paper cites.
Faster dynamic matrix inverse for faster lps
Jiang, S.; Song, Z.; Weinstein, O.; and Zhang, H. 2020c · 2004
Earlier work this paper cites.
LOCAL RADEMACHER COMPLEXITIES
Bartlett, P. L.; Bousquet, O.; and Mendelson, S. 2005 · 2005
Earlier work this paper cites.
Training (overparametrized) neural networks in near-linear time
Brand, J. v. d.; Peng, B.; Song, Z.; and Weinstein, O. 2020b · 2006
Earlier work this paper cites.
A direct formulation for sparse PCA using semidefinite programming
d’Aspremont, A.; Ghaoui, L. E.; Jordan, M. I.; and Lanckriet, G. R. G. 2006 · 2006
Earlier work this paper cites.
Robust Sub-Gaussian Principal Component Analysis and Width-Independent Schatten Packing
Jambulapati, A.; Li, J.; and Tian, K. 2020 · 2006
Earlier work this paper cites.
Training linear SVMs in linear time
Joachims, T. 2006 · 2006
Earlier work this paper cites.
The tradeoffs of large scale learning
Bottou, L.; and Bousquet, O. 2007 · 2007
Earlier work this paper cites.
High-dimensional analysis of semidefinite relaxations for sparse principal components
Amini, A. A.; and Wainwright, M. J. 2009 · 2009
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A.; Juditsky, A.; Lan, G.; and Shapiro, A. 2009 · 2009
Earlier work this paper cites.
Unifying Matrix Data Structures: Simplifying and Speeding up Iterative Algorithms
van den Brand, J. 2020 · 2010
Earlier work this paper cites.
Dong, S.; Lee, Y. T.; and Ye, G. 2023 · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E.; and Bach, F. 2011 · 2011
Earlier work this paper cites.
Agnostic learning of monomials by halfspaces is hard
Feldman, V.; Guruswami, V.; Raghavendra, P.; and Wu, Y. 2012 · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R.; and Zhang, T. 2013 · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Nesterov, Y. 2013 · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shalev-Shwartz, S.; and Zhang, T. 2013 · 2013
Earlier work this paper cites.
The nature of statistical learning theory
Vapnik, V. 2013 · 2013
Earlier work this paper cites.
Constant step size least-mean-square: Bias-variance trade-offs and optimal sampling distributions
Défossez, A.; and Bach, F. 2014 · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Bubeck, S.; et al. 2015 · 2015
Earlier work this paper cites.
Competing with the empirical risk minimizer in a single pass
Frostig, R.; Ge, R.; Kakade, S. M.; and Sidford, A. 2015 · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2016 · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Reed, S.; Akata, Z.; Yan, X.; Logeswaran, L.; Schiele, B.; and Lee, H. 2016 · 2016
Cited alongside, same era.
Improved techniques for training gans
Salimans, T.; Goodfellow, I.; Zaremba, W.; Cheung, V.; Radford, A.; and Chen, X. 2016 · 2016
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Cited alongside, same era.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Stochastic primal-dual coordinate method for regularized empirical risk minimization
Zhang, Y.; and Xiao, L. 2017 · 2017
A Faster Small Treewidth SDP Solver
Gu, Y.; and Song, Z. 2022 · 2022
Later among the works it cites.
Spvit: Enabling faster vision transformers via latency-aware soft token pruning
Kong, Z.; Dong, P.; Ma, X.; Meng, X.; Niu, W.; Sun, M.; Shen, X.; Yuan, G.; Ren, B.; Tang, H.; et al. 2022 · 2022
Later among the works it cites.
Autoregressive image generation using residual quantization
Lee, D.; Kim, C.; Kim, S.; Cho, M.; and Han, W.-S. 2022 · 2022
Later among the works it cites.
Li, Y.; Zhao, P.; Yuan, G.; Lin, X.; Wang, Y.; and Chen, X. 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S.; Zhai, X.; Poczos, B.; and Singh, A. 2018 · 2018
Cited alongside, same era.
On the local minima of the empirical risk
Jin, C.; Liu, L. T.; Ge, R.; and Jordan, M. I. 2018 · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y.; and Liang, Y. 2018 · 2018
Cited alongside, same era.
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Mihaylov, T.; Clark, P.; Khot, T.; and Sabharwal, A. 2018 · 2018
Cited alongside, same era.
Image transformer
Parmar, N.; et al. 2018 · 2018
Cited alongside, same era.
Compiler-aware neural architecture search for on-mobile real-time super-resolution
Wu, Y.; Gong, Y.; Zhao, P.; et al. 2022 · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J.; Xu, Y.; Koh, J. Y.; Luong, T.; Baid, G.; Wang, Z.; Vasudevan, V.; Ku, A.; Yang, Y.; Ayan, B. K.; et al. 2022 · 2022
Later among the works it cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Zhang, L. 2022 · 2022
Later among the works it cites.
Advancing model pruning via bi-level optimization
Zhang, Y.; Yao, Y.; Ram, P.; Zhao, P.; Chen, T.; Hong, M.; Wang, Y.; and Liu, S. 2022 · 2022
Later among the works it cites.
Achiam, J.; Adler, S.; et al. 2023 · 2023
Later among the works it cites.
Fluctuation-based Adaptive Structured Pruning for Large Language Models
An, Y.; Zhao, X.; Yu, T.; Tang, M.; and Wang, J. 2023 · 2023
Later among the works it cites.
Federated Empirical Risk Minimization via Second-Order Method
Bian, S.; Song, Z.; and Yin, J. 2023 · 2023
Later among the works it cites.
Algorithm and Hardness for Dynamic Attention Maintenance in Large Language Models
Brand, J. v. d.; Song, Z.; and Zhou, T. 2023 · 2023
Later among the works it cites.
Attention scheme inspired softmax regression
Deng, Y.; Li, Z.; and Song, Z. 2023 · 2023
Later among the works it cites.
Deng, Y.; Mahadevan, S.; and Song, Z. 2023 · 2023
Later among the works it cites.
Unmasking transformers: A theoretical approach to data recovery via attention weights
Deng, Y.; Song, Z.; Xie, S.; and Yang, C. 2023 · 2023
Later among the works it cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E.; and Alistarh, D. 2023 · 2023
Later among the works it cites.
An over-parameterized exponential regression
Gao, Y.; Mahadevan, S.; and Song, Z. 2023 · 2023
Later among the works it cites.
An iterative algorithm for rescaled hyperbolic functions regression
Gao, Y.; Song, Z.; and Yin, J. 2023 · 2023
Later among the works it cites.
Condense: A Framework for Device and Frequency Adaptive Neural Network Models on the Edge
Gong, Y.; Zhao, P.; Zhan, Z.; et al. 2023 · 2023
Later among the works it cites.
A nearly-linear time algorithm for structured support vector machines
Gu, Y.; Song, Z.; and Zhang, L. 2023 · 2023
Later among the works it cites.
Polysketchformer: Fast transformers via sketches for polynomial kernels
Kacham, P.; Mirrokni, V.; and Zhong, P. 2023 · 2023
Later among the works it cites.
Peeling the onion: Hierarchical reduction of data redundancy for efficient vision transformer training
Kong, Z.; Ma, H.; Yuan, G.; Sun, M.; Xie, Y.; Dong, P.; Meng, X.; Shen, X.; Tang, H.; Qin, M.; et al. 2023 · 2023
Later among the works it cites.
Solving regularized exp, cosh and sinh regression problems
Li, Z.; Song, Z.; and Zhou, T. 2023 · 2023
Later among the works it cites.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke, Q.; Song, Z.; Zhang, L.; and Zhuo, D. 2023 · 2023
Later among the works it cites.
Llm-pruner: On the structural pruning of large language models
Ma, X.; Fang, G.; and Wang, X. 2023 · 2023
Later among the works it cites.
Is Solving Graph Neural Tangent Kernel Equivalent to Training Graph Neural Network?
Qin, L.; Song, Z.; and Sun, B. 2023 · 2023
Later among the works it cites.
Efficient SGD Neural Network Training via Sublinear Activated Neuron Identification
Qin, L.; Song, Z.; and Yang, Y. 2023 · 2023
Later among the works it cites.
A Theoretical Analysis Of Nearest Neighbor Search On Approximate Near Neighbor Graph
Shrivastava, A.; Song, Z.; and Xu, Z. 2023 · 2023
Later among the works it cites.
A unified scheme of resnet and softmax
Song, Z.; Wang, W.; and Yin, J. 2023 · 2023
Later among the works it cites.
The Expressibility of Polynomial based Attention Scheme
Song, Z.; Xu, G.; and Yin, J. 2023 · 2023
Later among the works it cites.
Song, Z.; Ye, M.; and Zhang, L. 2023 · 2023
Later among the works it cites.
A Simple and Effective Pruning Approach for Large Language Models
Sun, M.; Liu, Z.; Bair, A.; and Kolter, J. Z. 2023 · 2023
Later among the works it cites.
Transformers as support vector machines
Tarzanagh, D. A.; Li, Y.; Thrampoulidis, C.; and Oymak, S. 2023 · 2023
Later among the works it cites.
How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation
Alman, J.; and Song, Z. 2024 · 2024
Closest in time.
Slicegpt: Compress large language models by deleting rows and columns
Ashkboos, S.; et al. 2024 · 2024
Closest in time.
How to Protect Copyright Data in Optimization of Large Language Models?
Chu, T.; Song, Z.; and Yang, C. 2024 · 2024
Closest in time.
Quantum Speedup for Spectral Approximation of Kronecker Products
Gao, Y.; Song, Z.; Zhang, R.; and Zhou, Y. 2024 · 2024
Closest in time.
Fast Second-order Method for Neural Network under Small Treewidth Setting
Li, X.; Long, J.; Song, Z.; and Zhou, T. 2024c · 2024
Closest in time.
LLaMA 3: The most capable openly available LLM to date
Meta. 2024 · 2024
Closest in time.
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Sun, P.; Jiang, Y.; Chen, S.; Zhang, S.; Peng, B.; Luo, P.; and Yuan, Z. 2024 · 2024
Closest in time.
Pruning Foundation Models for High Accuracy without Retraining
Zhao, P.; Sun, F.; Shen, X.; et al. 2024 · 2024
Closest in time.