Fetching the paper…
Reading the bibliography…
Diffusion Transformers have emerged as the preeminent models for a wide array of generative tasks, demonstrating superior performance and efficacy across various applications.
Improved Precision and Recall Metric for Assessing Generative Models
Kynkäänniemi, T.; Karras, T.; Laine, S.; Lehtinen, J.; and Aila, T. 2019 · 1904
Earlier work this paper cites.
BERT Rediscovers the Classical NLP Pipeline
Tenney, I.; Das, D.; and Pavlick, E. 2019 · 1905
Earlier work this paper cites.
What Does BERT Look At? An Analysis of BERT’s Attention
Clark, K.; Khandelwal, U.; Levy, O.; and Manning, C. D. 2019 · 1906
Earlier work this paper cites.
Quantum Entropy Scoring for Fast Robust Mean Estimation and Improved Outlier Detection
Dong, Y.; Hopkins, S. B.; and Li, J. 2019 · 1906
Earlier work this paper cites.
Analyzing the Structure of Attention in a Transformer Language Model
Vig, J.; and Belinkov, Y. 2019 · 1906
Earlier work this paper cites.
Reducing Transformer Depth on Demand with Structured Dropout
Fan, A.; Grave, E.; and Joulin, A. 2019 · 1909
Earlier work this paper cites.
Designing and Interpreting Probes with Control Tasks
Hewitt, J.; and Liang, P. 2019 · 1909
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O (1/k2)
Nesterov, Y. 1983 · 1983
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vapnik, V. 1991 · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T.; and Juditsky, A. B. 1992 · 1992
Earlier work this paper cites.
The Volumetric Barrier for Semidefinite Programming
Anstreicher, K. M. 2000 · 2000
Earlier work this paper cites.
Training v-support vector classifiers: theory and algorithms
Chang, C.-C.; and Lin, C.-J. 2001 · 2001
Earlier work this paper cites.
An Improved Cutting Plane Method for Convex Optimization, Convex-Concave Games and its Applications
Jiang, H.; Lee, Y. T.; Song, Z.; and wai Wong, S. C. 2020b · 2004
Earlier work this paper cites.
Faster dynamic matrix inverse for faster lps
Jiang, S.; Song, Z.; Weinstein, O.; and Zhang, H. 2020c · 2004
Earlier work this paper cites.
Local rademacher complexities
Bartlett, P. L.; Bousquet, O.; and Mendelson, S. 2005 · 2005
Earlier work this paper cites.
A direct formulation for sparse PCA using semidefinite programming
d’Aspremont, A.; Ghaoui, L. E.; Jordan, M. I.; and Lanckriet, G. R. G. 2006 · 2006
Earlier work this paper cites.
Robust Sub-Gaussian Principal Component Analysis and Width-Independent Schatten Packing
Jambulapati, A.; Li, J.; and Tian, K. 2020 · 2006
Earlier work this paper cites.
Training linear SVMs in linear time
Joachims, T. 2006 · 2006
Earlier work this paper cites.
The tradeoffs of large scale learning
Bottou, L.; and Bousquet, O. 2007 · 2007
Earlier work this paper cites.
High-dimensional analysis of semidefinite relaxations for sparse principal components
Amini, A. A.; and Wainwright, M. J. 2009 · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A.; Juditsky, A.; Lan, G.; and Shapiro, A. 2009 · 2009
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; et al. 2021 · 2010
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Unifying Matrix Data Structures: Simplifying and Speeding up Iterative Algorithms
van den Brand, J. 2020 · 2010
Earlier work this paper cites.
Dong, S.; Lee, Y. T.; and Ye, G. 2023 · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E.; and Bach, F. 2011 · 2011
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2020 · 2011
Earlier work this paper cites.
Agnostic learning of monomials by halfspaces is hard
Feldman, V.; Guruswami, V.; Raghavendra, P.; and Wu, Y. 2012 · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R.; and Zhang, T. 2013 · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Nesterov, Y. 2013 · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shalev-Shwartz, S.; and Zhang, T. 2013 · 2013
Earlier work this paper cites.
Introduction to real analysis
Trench, W. F. 2013 · 2013
Earlier work this paper cites.
The nature of statistical learning theory
Vapnik, V. 2013 · 2013
Earlier work this paper cites.
Constant step size least-mean-square: Bias-variance trade-offs and optimal sampling distributions
Défossez, A.; and Bach, F. 2014 · 2014
Earlier work this paper cites.
Competing with the empirical risk minimizer in a single pass
Frostig, R.; Ge, R.; Kakade, S. M.; and Sidford, A. 2015 · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Deep Unsupervised Learning using Nonequilibrium Thermodynamics
Sohl-Dickstein, J.; Weiss, E. A.; Maheswaranathan, N.; and Ganguli, S. 2015 · 2015
Earlier work this paper cites.
TensorFlow: A system for large-scale machine learning
Abadi, M.; Barham, P.; Chen, J.; Chen, Z.; et al. 2016 · 2016
Earlier work this paper cites.
Improved techniques for training gans
Salimans, T.; Goodfellow, I.; Zaremba, W.; Cheung, V.; Radford, A.; and Chen, X. 2016 · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Stochastic primal-dual coordinate method for regularized empirical risk minimization
Zhang, Y.; and Xiao, L. 2017 · 2017
Cited alongside, same era.
TVM: An automated end-to-end optimizing compiler for deep learning
Chen, T.; Moreau, T.; Jiang, Z.; et al. 2018 · 2018
Cited alongside, same era.
On the local minima of the empirical risk
Jin, C.; Liu, L. T.; Ge, R.; and Jordan, M. I. 2018 · 2018
Cited alongside, same era.
Memory-based Parameter Adaptation
Sprechmann, P.; Jayakumar, S. M.; Rae, J. W.; Pritzel, A.; Badia, A. P.; Uria, B.; Vinyals, O.; Hassabis, D.; Pascanu, R.; and Blundell, C. 2018 · 2018
Cited alongside, same era.
Robust Estimators in High Dimensions without the Computational Intractability
Unmasking transformers: A theoretical approach to data recovery via attention weights
Deng, Y.; Song, Z.; Xie, S.; and Yang, C. 2023 · 2023
Later among the works it cites.
Shortcut learning of large language models in natural language understanding
Du, M.; He, F.; Zou, N.; Tao, D.; and Hu, X. 2023 · 2023
Later among the works it cites.
Structural pruning for diffusion models
Fang, G.; Ma, X.; and Wang, X. 2023 · 2023
Later among the works it cites.
An over-parameterized exponential regression
Gao, Y.; Mahadevan, S.; and Song, Z. 2023 · 2023
Later among the works it cites.
An iterative algorithm for rescaled hyperbolic functions regression
Gao, Y.; Song, Z.; and Yin, J. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diakonikolas, I.; Kamath, G.; Kane, D.; Li, J.; Moitra, A.; and Stewart, A. 2019 · 2019
Cited alongside, same era.
Solving empirical risk minimization in the current matrix multiplication time
Lee, Y. T.; Song, Z.; and Zhang, Q. 2019 · 2019
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Song, Y.; and Ermon, S. 2019 · 2019
Cited alongside, same era.
A deterministic linear program solver in current matrix multiplication time
Brand, J. v. d. 2020 · 2020
Cited alongside, same era.
Solving tall dense linear programs in nearly linear time
Brand, J. v. d.; Lee, Y. T.; Sidford, A.; and Song, Z. 2020 · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Cited alongside, same era.
A faster interior point method for semidefinite programming
Jiang, H.; Kathuria, T.; Lee, Y. T.; Padmanabhan, S.; and Song, Z. 2020a · 2020
Cited alongside, same era.
A nearly-linear time algorithm for structured support vector machines
Gu, Y.; Song, Z.; and Zhang, L. 2023 · 2023
Later among the works it cites.
PTQD: Accurate Post-Training Quantization for Diffusion Models
He, Y.; Liu, L.; Liu, J.; Wu, W.; Zhou, H.; and Zhuang, B. 2023 · 2023
Later among the works it cites.
Measuring Forgetting of Memorized Training Examples
Jagielski, M.; Thakkar, O.; Tramèr, F.; Ippolito, D.; Lee, K.; Carlini, N.; Wallace, E.; Song, S.; Thakurta, A.; Papernot, N.; and Zhang, C. 2023 · 2023
Later among the works it cites.
Polysketchformer: Fast transformers via sketches for polynomial kernels
Kacham, P.; Mirrokni, V.; and Zhong, P. 2023 · 2023
Later among the works it cites.
BK-SDM: Architecturally Compressed Stable Diffusion for Efficient Text-to-Image Generation
Kim, B.-K.; Song, H.-K.; Castells, T.; and Choi, S. 2023 · 2023
Later among the works it cites.
Peeling the onion: Hierarchical reduction of data redundancy for efficient vision transformer training
Kong, Z.; Ma, H.; Yuan, G.; et al. 2023 · 2023
Later among the works it cites.
Solving regularized exp, cosh and sinh regression problems
Li, Z.; Song, Z.; and Zhou, T. 2023 · 2023
Later among the works it cites.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke, Q.; Song, Z.; Zhang, L.; and Zhuo, D. 2023 · 2023
Later among the works it cites.
Vdt: General-purpose video diffusion transformers via mask modeling
Lu, H.; Yang, G.; Fei, N.; Huo, Y.; Lu, Z.; Luo, P.; and Ding, M. 2023 · 2023
Later among the works it cites.
Scalable diffusion models with transformers
Peebles, W.; and Xie, S. 2023 · 2023
Later among the works it cites.
Is Solving Graph Neural Tangent Kernel Equivalent to Training Graph Neural Network?
Qin, L.; Song, Z.; and Sun, B. 2023 · 2023
Later among the works it cites.
A Theoretical Analysis Of Nearest Neighbor Search On Approximate Near Neighbor Graph
Shrivastava, A.; Song, Z.; and Xu, Z. 2023 · 2023
Later among the works it cites.
A unified scheme of resnet and softmax
Song, Z.; Wang, W.; and Yin, J. 2023 · 2023
Later among the works it cites.
The Expressibility of Polynomial based Attention Scheme
Song, Z.; Xu, G.; and Yin, J. 2023 · 2023
Later among the works it cites.
Song, Z.; Ye, M.; and Zhang, L. 2023 · 2023
Later among the works it cites.
Transformers as support vector machines
Tarzanagh, D. A.; Li, Y.; Thrampoulidis, C.; and Oymak, S. 2023 · 2023
Later among the works it cites.
Cache Me if You Can: Accelerating Diffusion Models through Block Caching
Wimbauer, F.; Wu, B.; Schoenfeld, E.; et al. 2023 · 2023
Later among the works it cites.
LLaMA-Adapter: Efficient Finetuning of Language Models with Zero-init Attention
Zhang, R.; et al. 2023 · 2023
Later among the works it cites.
How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation
Alman, J.; and Song, Z. 2024 · 2024
Closest in time.
LD-Pruner: Efficient Pruning of Latent Diffusion Models using Task-Agnostic Insights
Castells, T.; Song, H.-K.; Kim, B.-K.; and Choi, S. 2024 · 2024
Closest in time.
Quantum Speedup for Spectral Approximation of Kronecker Products
Gao, Y.; Song, Z.; Zhang, R.; and Zhou, Y. 2024 · 2024
Closest in time.
Open-Sora-Plan
Lab, P.-Y.; and etc., T. A. 2024 · 2024
Closest in time.
Fast Second-order Method for Neural Network under Small Treewidth Setting
Li, X.; Long, J.; Song, Z.; and Zhou, T. 2024c · 2024
Closest in time.
SDXL-Lightning: Progressive Adversarial Diffusion Distillation
Lin, S.; Wang, A.; and Yang, X. 2024 · 2024
Closest in time.
Latent Consistency Models: Synthesizing High-Resolution Images with Few-step Inference
Luo, S.; Tan, Y.; Huang, L.; Li, J.; and Zhao, H. 2024 · 2024
Closest in time.
Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching
Ma, X.; Fang, G.; Mi, M. B.; and Wang, X. 2024 · 2024
Closest in time.
DeepCache: Accelerating Diffusion Models for Free
Ma, X.; Fang, G.; and Wang, X. 2024 · 2024
Closest in time.
Video generation models as world simulators
OpenAI. 2024 · 2024
Closest in time.
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
Raposo, D.; Ritter, S.; Richards, B.; et al. 2024 · 2024
Closest in time.
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Sun, P.; Jiang, Y.; Chen, S.; Zhang, S.; Peng, B.; Luo, P.; and Yuan, Z. 2024 · 2024
Closest in time.
One-step Diffusion with Distribution Matching Distillation
Yin, T.; Gharbi, M.; Zhang, R.; Shechtman, E.; Durand, F.; Freeman, W. T.; and Park, T. 2024 · 2024
Closest in time.
LAPTOP-Diff: Layer Pruning and Normalized Distillation for Compressing Diffusion Models
Zhang, D.; Li, S.; Chen, C.; Xie, Q.; and Lu, H. 2024 · 2024
Closest in time.
Pruning Foundation Models for High Accuracy without Retraining
Zhao, P.; Sun, F.; Shen, X.; Yu, P.; Kong, Z.; Wang, Y.; and Lin, X. 2024 · 2024
Closest in time.
Open-Sora: Democratizing Efficient Video Production for All
Zheng, Z.; Peng, X.; Yang, T.; Shen, C.; Li, S.; Liu, H.; Zhou, Y.; Li, T.; and You, Y. 2024 · 2024
Closest in time.