Fetching the paper…
Reading the bibliography…
This paper presents Diffusion Forcing, a new training paradigm where a diffusion model is trained to denoise a set of tokens with independent per-token noise levels.
Scoring rules for continuous probability distributions
J. E. Matheson and R. L. Winkler · 1976
Earlier work this paper cites.
Model predictive control: Theory and practice—a survey
C. E. Garcia, D. M. Prett, and M. Morari · 1989
Earlier work this paper cites.
A Learning Algorithm for Continually Running Fully Recurrent Neural Networks
R. J. Williams and D. Zipser · 1989
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Learning to forget: continual prediction with lstm
F. Gers, J. Schmidhuber, and F. Cummins · 1999
Earlier work this paper cites.
Go-garch: A multivariate generalized orthogonal garch model
R. van der Weide · 2002
Earlier work this paper cites.
New Introduction to Multiple Time Series Analysis
H. Lütkepohl · 2005
Earlier work this paper cites.
Forecasting with Exponential Smoothing: The State Space Approach
R. Hyndman, A. B. Koehler, J. K. Ord, and R. D. Snyder · 2008
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, Ç. Gülçehre, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
A recurrent latent variable model for sequential data
J. Chung, K. Kastner, L. Dinh, K. Goel, A. C. Courville, and Y. Bengio · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
G. Williams, A. Aldrich, and E. Theodorou · 2015
Earlier work this paper cites.
Temporal regularized matrix factorization
H. Yu, N. Rao, and I. S. Dhillon · 2015
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves, et al · 2016
Earlier work this paper cites.
Structured inference networks for nonlinear state space models
R. G. Krishnan, U. Shalit, and D. Sontag · 2017
Earlier work this paper cites.
Modeling long- and short-term temporal patterns with deep neural networks
G. Lai, W. Chang, Y. Yang, and H. Liu · 2017
Earlier work this paper cites.
Dart: Noise injection for robust imitation learning
M. Laskey, J. Lee, R. Fox, A. Dragan, and K. Goldberg · 2017
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Learning latent dynamics for planning from pixels
D. Hafner, T. P. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2018
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. de Oliveira Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
J. Cohen, E. Rosenfeld, and Z. Kolter · 2019
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Earlier work this paper cites.
High-dimensional multivariate forecasting with low-rank gaussian copula processes
D. Salinas, M. Bohlke-Schneider, L. Callot, R. Medico, J. Gasthaus, and R. Medico · 2019
Cited alongside, same era.
Gluonts: Probabilistic and neural time series modeling in python
A. Alexandrov, K. Benidis, M. Bohlke-Schneider, V. Flunkert, J. Gasthaus, T. Januschowski, D. C. Maddix, S. Rangapuram, D. Salinas, J. Schulz, L. Stella, A. C. Türkmen, and Y. Wang · 2020
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
Generative pretraining from pixels
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever · 2020
Cited alongside, same era.
Normalizing kalman filters for multivariate time series analysis
E. de Bézenac, S. S. Rangapuram, K. Benidis, M. Bohlke-Schneider, R. Kurle, L. Stella, H. Hasson, P. Gallinari, and T. Januschowski · 2020
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Later among the works it cites.
Classifier-free diffusion guidance, 2022
J. Ho and T. Salimans · 2022
Later among the works it cites.
Video diffusion models, 2022
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet · 2022
Later among the works it cites.
Planning with diffusion for flexible behavior synthesis
M. Janner, Y. Du, J. B. Tenenbaum, and S. Levine · 2022
Later among the works it cites.
Diffusion-lm improves controllable text generation, 2022
X. L. Li, J. Thickstun, I. Gulrajani, P. Liang, and T. B. Hashimoto · 2022
Later among the works it cites.
Progressive distillation for fast sampling of diffusion models
T. Salimans and J. Ho · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
D4RL: datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Cited alongside, same era.
Deepar: Probabilistic forecasting with autoregressive recurrent networks
D. Salinas, V. Flunkert, J. Gasthaus, and T. Januschowski · 2020
Cited alongside, same era.
Deepar: Probabilistic forecasting with autoregressive recurrent networks
D. Salinas, V. Flunkert, J. Gasthaus, and T. Januschowski · 2020
Cited alongside, same era.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Cited alongside, same era.
Later among the works it cites.
Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, V. Jampani, and R. Rombach · 2023
Later among the works it cites.
Provable guarantees for generative behavior cloning: Bridging low-level stability and high-level behavior
A. Block, A. Jadbabaie, D. Pfrommer, M. Simchowitz, and R. Tedrake · 2023
Later among the works it cites.
On the importance of noise scheduling for diffusion models, 2023
T. Chen · 2023
Later among the works it cites.
Masked diffusion transformer is a strong image synthesizer
S. Gao, P. Zhou, M.-M. Cheng, and S. Yan · 2023
Later among the works it cites.
Gaia-1: A generative world model for autonomous driving
A. Hu, L. Russell, H. Yeo, Z. Murez, G. Fedoseev, A. Kendall, J. Shotton, and G. Corrado · 2023
Later among the works it cites.
Diffusion-based generation, optimization, and planning in 3d scenes
S. Huang, Z. Wang, P. Li, B. Jia, T. Liu, Y. Zhu, W. Liang, and S.-C. Zhu · 2023
Later among the works it cites.
RWKV: Reinventing RNNs for the transformer era
B. Peng, E. Alcaide, Q. Anthony, A. Albalak, S. Arcadinho, S. Biderman, H. Cao, X. Cheng, M. Chung, L. Derczynski, X. Du, M. Grella, K. Gv, X. He, H. Hou, P. Kazienko, J. Kocon, J. Kong, B. Koptyra, H. Lau, J. Lin, K. S. I. Mantri, F. Mom, A. Saito, G. Song, X. Tang, J. Wind, S. Woźniak, Z. Zhang, Q. Zhou, J. Zhu, and R.-J. Zhu · 2023
Later among the works it cites.
Diffusion models as masked autoencoders
C. Wei, K. Mangalam, P.-Y. Huang, Y. Li, H. Fan, H. Xu, H. Wang, C. Xie, A. Yuille, and C. Feichtenhofer · 2023
Later among the works it cites.
Ar-diffusion: Auto-regressive diffusion model for text generation, 2023
T. Wu, Z. Fan, X. Liu, Y. Gong, Y. Shen, J. Jiao, H.-T. Zheng, J. Li, Z. Wei, J. Guo, N. Duan, and W. Chen · 2023
Later among the works it cites.
Ar-diffusion: Auto-regressive diffusion model for text generation
T. Wu, Z. Fan, X. Liu, H.-T. Zheng, Y. Gong, J. Jiao, J. Li, J. Guo, N. Duan, W. Chen, et al · 2023
Later among the works it cites.
Temporally consistent transformers for video generation, 2023
W. Yan, D. Hafner, S. James, and P. Abbeel · 2023
Later among the works it cites.
Diffusion probabilistic modeling for video generation
R. Yang, P. Srivastava, and S. Mandt · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models, 2023
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan · 2023
Later among the works it cites.
Tutorial on diffusion models for imaging and vision
S. H. Chan · 2024
Closest in time.
Diffusion policy: Visuomotor policy learning via action diffusion, 2024
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2024
Closest in time.
Efficient diffusion training via min-snr weighting strategy, 2024
T. Hang, S. Gu, C. Li, J. Bao, D. Chen, H. Hu, X. Geng, and B. Guo · 2024
Closest in time.
Understanding diffusion objectives as the elbo with simple data augmentation
D. Kingma and R. Gao · 2024
Closest in time.
D. Ruhe, J. Heek, T. Salimans, and E. Hoogeboom · 2024
Closest in time.