Fetching the paper…
Reading the bibliography…
The success of large generative models has driven a paradigm shift, leveraging massive multi-source data to enhance model capabilities.
Probability inequalities for likelihood ratios and convergence rates of sieve mles
Wong, W. H. and Shen, X · 1995
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
A model of inductive bias learning
Baxter, J · 2000
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
Geer, S. A · 2000
Earlier work this paper cites.
Empirical processes in statistics: Methods, examples, further problems, 2002
Wellner, J. A · 2002
Earlier work this paper cites.
A tutorial on energy-based learning
LeCun, Y., Chopra, S., Hadsell, R., Ranzato, M., Huang, F., et al · 2006
Earlier work this paper cites.
A notion of task relatedness yielding provable multiple-task learning guarantees
Ben-David, S. and Borbely, R. S · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Imagenet summary and statistics, 2010
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2010
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
The benefit of multitask representation learning
Maurer, A., Pontil, M., and Romera-Paredes, B · 2016
Earlier work this paper cites.
Neural autoregressive distribution estimation
Uria, B., Côté, M., Gregor, K., Murray, I., and Larochelle, H · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Imagenet hierarchy, 2018
Bostock., M · 2018
Earlier work this paper cites.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
Bartlett, P. L., Harvey, N., Liaw, C., and Mehrabian, A · 2019
Earlier work this paper cites.
Implicit generation and modeling with energy based models
Du, Y. and Mordatch, I · 2019
Earlier work this paper cites.
Generalization bounds for convolutional neural networks
Lin, S. and Zhang, J · 2019
Cited alongside, same era.
How multilingual is multilingual bert?
Pires, T., Schlinger, E., and Garrette, D · 2019
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Cited alongside, same era.
Adaptivity of deep relu network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality
Suzuki, T · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Diffusion models are minimax optimal distribution estimators
Oko, K., Akiyama, S., and Suzuki, T · 2023
Later among the works it cites.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
Toward understanding generative data augmentation
Zheng, C., Wu, G., and Li, C · 2023
Later among the works it cites.
Physics of language models: Part 3.1, knowledge storage and extraction
Allen-Zhu, Z. and Li, Y · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
On the theory of transfer learning: The importance of task diversity
Tripuraneni, N., Jordan, M. I., and Jin, C · 2020
Cited alongside, same era.
Norm-based generalisation bounds for deep multi-class convolutional neural networks
Ledent, A., Mustafa, W., Lei, Y., and Kloft, M · 2021
Cited alongside, same era.
Non-asymptotic excess risk bounds for classification with deep convolutional neural networks
Shen, G., Jiao, Y., Lin, Y., and Huang, J · 2021
Cited alongside, same era.
Towards understanding the data dependency of mixup-style training
Chidambaram, M., Wang, X., Hu, Y., Wu, C., and Ge, R · 2022
Cited alongside, same era.
Information-theoretic characterization of the generalization error for iterative semi-supervised learning
He, H., Yan, H., and Tan, V. Y · 2022
Cited alongside, same era.
Universality of empirical risk minimization
Montanari, A. and Saeed, B. N · 2022
Cited alongside, same era.
Pixart- α \alpha : Fast training of diffusion transformer for photorealistic text-to-image synthesis
Chen, J., Yu, J., Ge, C., Yao, L., Xie, E., Wang, Z., Kwok, J. T., Luo, P., Lu, H., and Li, Z · 2024
Later among the works it cites.
Universality laws for gaussian mixtures in generalized linear models
Dandi, Y., Stephan, L., Krzakala, F., Loureiro, B., and Zdeborová, L · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., Podell, D., Dockhorn, T., English, Z., and Rombach, R · 2024
Later among the works it cites.
Unveil conditional diffusion models with classifier-free guidance: A sharp statistical theory
Fu, H., Yang, Z., Wang, M., and Chen, M · 2024
Later among the works it cites.
On the provable advantage of unsupervised pretraining
Ge, J., Tang, S., Fan, J., and Jin, C · 2024
Later among the works it cites.
Convergence analysis of flow matching in latent space with transformers
Jiao, Y., Lai, Y., Wang, Y., and Yan, B · 2024
Later among the works it cites.
Analyzing and improving the training dynamics of diffusion models
Karras, T., Aittala, M., Lehtinen, J., Hellsten, J., Aila, T., and Laine, S · 2024
Later among the works it cites.
Video generation models as world simulators
OpenAI · 2024
Later among the works it cites.
Ou, W. and Bölcskei, H · 2024
Later among the works it cites.
Exploring the complexity of deep neural networks through functional equivalence
Shen, G · 2024
Later among the works it cites.
Sequence length independent norm-based generalization bounds for transformers
Trauger, J. and Tewari, A · 2024
Later among the works it cites.
Metadata conditioning accelerates language model pre-training
Gao, T., Wettig, A., He, L., Dong, Y., Malladi, S., and Chen, D · 2025
Closest in time.