Fetching the paper…
Reading the bibliography…
Autonomous driving promises transformative improvements to transportation, but building systems capable of safely navigating the unstructured complexity of real-world scenarios remains challenging.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
J. Johnson, A. Alahi, and L. Fei-Fei · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
A. van den Oord, N. Kalchbrenner, L. Espeholt, k. kavukcuoglu, O. Vinyals, and A. Graves · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Earlier work this paper cites.
NIPS 2016 Tutorial: Generative Adversarial Networks
I. Goodfellow · 2016
Earlier work this paper cites.
Neural discrete representation learning
A. van den Oord, O. Vinyals, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Video Pixel Networks
N. Kalchbrenner, A. van den Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
D. Ha and J. Schmidhuber · 2018
Earlier work this paper cites.
Stochastic variational video prediction
M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine · 2018
Earlier work this paper cites.
Stochastic Video Generation with a Learned Prior
E. Denton and R. Fergus · 2018
Earlier work this paper cites.
MoCoGAN: Decomposing motion and content for video generation
S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz · 2018
Earlier work this paper cites.
Learning to drive in a day
A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J.-M. Allen, V.-D. Lam, A. Bewley, and A. Shah · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
High fidelity video prediction with large stochastic recurrent neural networks
R. Villegas, A. Pathak, H. Kannan, D. Erhan, Q. Le, and H. Lee · 2019
Earlier work this paper cites.
Adversarial Video Generation on Complex Datasets
A. Clark, J. Donahue, and K. Simonyan · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
nuScenes: A multimodal dataset for autonomous driving
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom · 2020
Earlier work this paper cites.
Probabilistic Future Prediction for Video Scene Understanding
A. Hu, F. Cotter, N. Mohan, C. Gurau, and A. Kendall · 2020
Earlier work this paper cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. Lillicrap, and D. Silver · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Stochastic latent residual video prediction
J.-Y. Franceschi, E. Delasalles, M. Chen, S. Lamprier, and P. Gallinari · 2020
Earlier work this paper cites.
Scaling autoregressive video models
D. Weissenborn, O. Täckström, and J. Uszkoreit · 2020
Earlier work this paper cites.
Data-efficient reinforcement learning with self-predictive representations
M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. C. Courville, and P. Bachman · 2020
Earlier work this paper cites.
Model-based reinforcement learning for atari
L. Kaiser, M. Babaeizadeh, P. Milos, B. Osinski, R. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, A. Mohiuddin, R. Sepassi, G. Tucker, and H. Michalewski · 2020
Cited alongside, same era.
Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset
S. Ettinger, S. Cheng, B. Caine, C. Liu, H. Zhao, S. Pradhan, Y. Chai, B. Sapp, C. Qi, Y. Zhou, Z. Yang, A. Chouard, P. Sun, J. Ngiam, V. Vasudevan, A. McCauley, J. Shlens, and D. Anguelov · 2021
Cited alongside, same era.
FIERY: Future Instance Prediction in Bird’s-Eye View From Surround Monocular Cameras
A. Hu, Z. Murez, N. Mohan, S. Dudas, J. Hawke, V. Badrinarayanan, R. Cipolla, and A. Kendall · 2021
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
M. Janner, Q. Li, and S. Levine · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Cited alongside, same era.
StyleGAN-V: A continuous video generator with the price, image quality and perks of StyleGAN2
I. Skorokhodov, S. Tulyakov, and M. Elhoseiny · 2022
Later among the works it cites.
Generating long videos of dynamic scenes
T. Brooks, J. Hellsten, M. Aittala, T.-C. Wang, T. Aila, J. Lehtinen, M.-Y. Liu, A. Efros, and T. Karras · 2022
Later among the works it cites.
MCVD: Masked conditional video diffusion for prediction, generation, and interpolation
V. Voleti, A. Jolicoeur-Martineau, and C. Pal · 2022
Later among the works it cites.
Diffusion models for video prediction and infilling
T. Höppe, A. Mehrjou, S. Bauer, D. Nielsen, and A. Dittadi · 2022
Later among the works it cites.
Make-A-Video: Text-to-video generation without text-video data
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y. Taigman · 2022
Later among the works it cites.
Magicvideo: Efficient video generation with latent diffusion models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Cited alongside, same era.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2021
Cited alongside, same era.
Fitvid: Overfitting in pixel-level video prediction
M. Babaeizadeh, M. Saffar, S. Nair, S. Levine, C. Finn, and D. Erhan · 2021
Cited alongside, same era.
DriveGAN: Towards a controllable high-quality neural simulation
S. W. Kim, J. Philion, A. Torralba, and S. Fidler · 2021
Cited alongside, same era.
VideoGPT: Video generation using vq-vae and transformers
W. Yan, Y. Zhang, P. Abbeel, and A. Srinivas · 2021
Cited alongside, same era.
CCVS: Context-aware controllable video synthesis
G. L. Moing, J. Ponce, and C. Schmid · 2021
Cited alongside, same era.
D. Zhou, W. Wang, H. Yan, W. Lv, Y. Zhu, and J. Feng · 2022
Later among the works it cites.
Long video generation with time-agnostic vqgan and time-sensitive transformer
S. Ge, T. Hayes, H. Yang, X. Yin, G. Pang, D. Jacobs, J.-B. Huang, and D. Parikh · 2022
Later among the works it cites.
HARP: Autoregressive latent video prediction with high-fidelity image generator
Y. Seo, K. Lee, F. Liu, S. James, and P. Abbeel · 2022
Later among the works it cites.
General-purpose, long-context autoregressive modeling with Perceiver AR
C. Hawthorne, A. Jaegle, C. Cangea, S. Borgeaud, C. Nash, M. Malinowski, S. Dieleman, O. Vinyals, M. Botvinick, I. Simon, H. Sheahan, N. Zeghidour, J.-B. Alayrac, J. Carreira, and J. Engel · 2022
Later among the works it cites.
Masked autoencoding for scalable and generalizable decision making
F. Liu, H. Liu, A. Grover, and P. Abbeel · 2022
Later among the works it cites.
GLaM: Efficient scaling of language models with mixture-of-experts
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. Bosma, Z. Zhou, T. Wang, Y. E. Wang, K. Webster, M. Pellat, K. Robinson, K. S. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. V. Le, Y. Wu, Z. Chen, and C. Cui · 2022
Later among the works it cites.
Scale efficiently: Insights from pre-training and fine-tuning transformers
Y. Tay, M. Dehghani, J. Rao, W. Fedus, S. Abnar, H. W. Chung, S. Narang, D. Yogatama, A. Vaswani, and D. Metzler · 2022
Later among the works it cites.
Planning-oriented autonomous driving
Y. Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wang, L. Lu, X. Jia, Q. Liu, J. Dai, Y. Qiao, and H. Li · 2023
Closest in time.
Transformers are sample-efficient world models
V. Micheli, E. Alonso, and F. Fleuret · 2023
Closest in time.
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Closest in time.
Daydreamer: World models for physical robot learning
P. Wu, A. Escontrela, D. Hafner, P. Abbeel, and K. Goldberg · 2023
Closest in time.
Structure and content-guided video synthesis with diffusion models
P. Esser, J. Chiu, P. Atighehchian, J. Granskog, and A. Germanidis · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Closest in time.
GPT-4 Technical Report
OpenAI · 2023
Closest in time.
simple diffusion: End-to-end diffusion for high resolution images
E. Hoogeboom, J. Heek, and T. Salimans · 2023
Closest in time.
Muse: Text-to-image generation via masked generative transformers
H. Chang, H. Zhang, J. Barber, A. Maschinot, J. Lezama, L. Jiang, M.-H. Yang, K. Murphy, W. T. Freeman, M. Rubinstein, et al · 2023
Closest in time.
Dreamix: Video diffusion models are general video editors
E. Molad, E. Horwitz, D. Valevski, A. R. Acha, Y. Matias, Y. Pritch, Y. Leviathan, and Y. Hoshen · 2023
Closest in time.
Align your latents: High-resolution video synthesis with latent diffusion models
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis · 2023
Closest in time.
Temporally consistent transformers for video generation
W. Yan, D. Hafner, S. James, and P. Abbeel · 2023
Closest in time.
Phenaki: Variable length video generation from open domain textual description
R. Villegas, M. Babaeizadeh, P.-J. Kindermans, H. Moraldo, H. Zhang, M. T. Saffar, S. Castro, J. Kunze, and D. Erhan · 2023
Closest in time.
MAGVIT: Masked Generative Video Transformer
L. Yu, Y. Cheng, K. Sohn, J. Lezama, H. Zhang, H. Chang, A. G. Hauptmann, M.-H. Yang, Y. Hao, I. Essa, and L. Jiang · 2023
Closest in time.
Masked trajectory models for prediction, representation, and control
P. Wu, A. Majumdar, K. Stone, Y. Lin, I. Mordatch, P. Abbeel, and A. Rajeswaran · 2023
Closest in time.
DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving
X. Wang, Z. Zhu, G. Huang, X. Chen, and J. Lu · 2023
Closest in time.
https://github.com/commaai/commavq , 2023
commaVQ · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, R. Jenatton, L. Beyer, M. Tschannen, A. Arnab, X. Wang, C. Riquelme, M. Minderer, J. Puigcerver, U. Evci, M. Kumar, S. van Steenkiste, G. F. Elsayed, A. Mahendran, F. Yu, A. Oliver, F. Huot, J. Bastings, M. P. Collier, A. Gritsenko, V. Birodkar, C. Vasconcelos, Y. Tay, T. Mensink, A. Kolesnikov, F. Pavetić, D. Tran, T. Kipf, M. Lučić, X. Zhai, D. Keysers, J. Harmsen, and N. Houlsby · 2023
Closest in time.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Closest in time.