Fetching the paper…
Reading the bibliography…
Large monolithic generative models trained on massive amounts of data have become an increasingly dominant approach in AI research.
Pandemonium: A paradigm for learning
Selfridge, O. G · 1988
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Hinton, G. E · 2002
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Koller, D. and Friedman, N · 2009
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Improved contrastive divergence training of energy based models
Du, Y., Li, S., Tenenbaum, J., and Mordatch, I · 2012
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Pixel recurrent neural networks
Van Den Oord, A., Kalchbrenner, N., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Implicit generation and generalization in energy-based models
Du, Y. and Mordatch, I · 2019
Earlier work this paper cites.
Your classifier is secretly an energy based model and you should treat it like one
Grathwohl, W., Wang, K.-C., Jacobsen, J.-H., Duvenaud, D., Norouzi, M., and Swersky, K · 2019
Earlier work this paper cites.
On the anatomy of mcmc-based maximum likelihood learning of energy-based models
Nijkamp, E., Hill, M., Han, T., Zhu, S.-C., and Wu, Y. N · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
A short note on learning discrete distributions
Canonne, C. L · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Unsupervised learning of compositional energy concepts
Du, Y., Li, S., Sharma, Y., Tenenbaum, B. J., and Mordatch, I · 2021
Earlier work this paper cites.
Oops i took a gradient: Scalable sampling for discrete distributions
Grathwohl, W., Swersky, K., Hashemi, M., Duvenaud, D., and Maddison, C · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Understanding the capabilities, limitations, and societal impact of large language models
Tamkin, A., Brundage, M., Clark, J., and Ganguli, D · 2021
Cited alongside, same era.
Is conditional generative modeling all you need for decision-making?
Ajay, A., Du, Y., Gupta, A., Tenenbaum, J., Jaakkola, T., and Agrawal, P · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jian, L., Lin, B. Y., West, P., Bhagavatula, C., Bras, R. L., Hwang, J. D., et al · 2023
Later among the works it cites.
Compositional sculpting of iterative generative processes
Garipov, T., De Peuter, S., Yang, G., Garg, V., Kaski, S., and Jaakkola, T · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Later among the works it cites.
Unsupervised compositional concepts discovery with text-to-image generative models
Liu, N., Du, Y., Li, S., Tenenbaum, J. B., and Torralba, A · 2023
Later among the works it cites.
Are emergent abilities in large language models just in-context learning?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J. B., and Levine, S · 2022
Cited alongside, same era.
Composing ensembles of pre-trained models via iterative consensus
Li, S., Du, Y., Tenenbaum, J. B., Torralba, A., and Mordatch, I · 2022
Cited alongside, same era.
Compositional visual generation with composable diffusion models
Liu, N., Li, S., Du, Y., Torralba, A., and Tenenbaum, J. B · 2022
Cited alongside, same era.
Vael: Bridging variational autoencoders and probabilistic logic programming
Misino, E., Marra, G., and Sansone, E · 2022
Cited alongside, same era.
Probabilistic machine learning: an introduction
Murphy, K. P · 2022
Cited alongside, same era.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2022
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Lu, S., Bigoulaeva, I., Sachdeva, R., Madabushi, H. T., and Gurevych, I · 2023
Later among the works it cites.
Openeqa: Embodied question answering in the era of foundation models
Majumdar, A., Ajay, A., Zhang, X., Putta, P., Yenamandra, S., Henaff, M., Silwal, S., Mcvay, P., Maksymets, O., Arnaud, S., Yadav, K., Li, Q., Newman, B., Sharma, M., Berges, V., Zhang, S., Agrawal, P., Bisk, Y., Batra, D., Kalakrishnan, M., Meier, F., Paxton, C., Sax, S., and Rajeswaran, A · 2023
Later among the works it cites.
Generative skill chaining: Long-horizon skill planning with diffusion models
Mishra, U. A., Xue, S., Chen, Y., and Xu, D · 2023
Later among the works it cites.
Are emergent abilities of large language models a mirage?
Schaeffer, R., Miranda, B., and Koyejo, S · 2023
Later among the works it cites.
Neurosymbolic grounding for compositional world models
Sehgal, A., Grayeli, A., Sun, J. J., and Chaudhuri, S · 2023
Later among the works it cites.
Pretraining data mixtures enable narrow model selection capabilities in transformer models
Yadlowsky, S., Doshi, L., and Tripuraneni, N · 2023
Later among the works it cites.
Key trends and figures in machine learning, 2023
Epoch · 2024
Closest in time.
Additive decoders for latent variables identification and cartesian-product extrapolation
Lachapelle, S., Mahajan, D., Mitliagkas, I., and Lacoste-Julien, S · 2024
Closest in time.
Compositional image decomposition with diffusion models, 2024
Su, J., Liu, N., Tenenbaum, J. B., and Du, Y · 2024
Closest in time.
Poco: Policy composition from and for heterogeneous robot learning
Wang, L., Zhao, J., Du, Y., Adelson, E. H., and Tedrake, R · 2024
Closest in time.
Compositional generalization from first principles
Wiedemer, T., Mayilvahanan, P., Bethge, M., and Brendel, W · 2024
Closest in time.