Fetching the paper…
Reading the bibliography…
In this paper, we investigate the challenge of spatio-temporal video prediction task, which involves generating future video frames based on historical spatio-temporal observation streams.
Maximum likelihood estimation of intrinsic dimension. In Proceedings of the Conference on Neural Information Processing Systems
Elizaveta Levina and Peter Bickel. 2004 · 2004
Earlier work this paper cites.
Recognizing human actions: a local SVM approach. In Proceedings of the 17th International Conference on Pattern Recognition, 2004. ICPR 2004. , Vol. 3. 32–36 Vol.3
C. Schuldt, I. Laptev, and B. Caputo. 2004 · 2004
Earlier work this paper cites.
Convolutional LSTM network: A machine learning approach for precipitation nowcasting. In Proceedings of the Conference on Neural Information Processing Systems
Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo. 2015 · 2015
Earlier work this paper cites.
Unsupervised Learning of Video Representations Using LSTMs. In Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37 (Lille, France) (ICML’15) . JMLR.org, 843–852
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhutdinov. 2015 · 2015
Earlier work this paper cites.
A guide to convolution arithmetic for deep learning
Vincent Dumoulin and Francesco Visin. 2016 · 2016
Earlier work this paper cites.
Stochastic variational video prediction. In Proceedings of the International Conference on Learning Representations
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H Campbell, and Sergey Levine. 2017 · 2017
Earlier work this paper cites.
Geometric deep learning: going beyond euclidean data
Michael M Bronstein, Joan Bruna, Yann LeCun, Arthur Szlam, and Pierre Vandergheynst. 2017 · 2017
Earlier work this paper cites.
Hashnet: Deep learning to hash by continuation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 5608–5617
Zhangjie Cao, Mingsheng Long, Jianmin Wang, and Philip S Yu. 2017 · 2017
Earlier work this paper cites.
Video frame synthesis using deep voxel flow. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4463–4471
Ziwei Liu, Raymond A Yeh, Xiaoou Tang, Yiming Liu, and Aseem Agarwala. 2017 · 2017
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning. In Proceedings of the International Conference on Learning Representations
William Lotter, Gabriel Kreiman, and David Cox. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning. In Proceedings of the Conference on Neural Information Processing Systems
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Decomposing motion and content for natural video sequence prediction. In Proceedings of the International Conference on Learning Representations
Ruben Villegas, Jimei Yang, Seunghoon Hong, Xunyu Lin, and Honglak Lee. 2017 · 2017
Earlier work this paper cites.
Deep spatio-temporal residual networks for citywide crowd flows prediction. In Proceedings of the AAAI conference on artificial intelligence , Vol. 31
Junbo Zhang, Yu Zheng, and Dekang Qi. 2017 · 2017
Earlier work this paper cites.
Pde-net: Learning pdes from data. In Proceedings of the International Conference on Machine Learning . 3208–3216
Zichao Long, Yiping Lu, Xianzhong Ma, and Bin Dong. 2018 · 2018
Earlier work this paper cites.
Deep hidden physics models: Deep learning of nonlinear partial differential equations
Maziar Raissi. 2018 · 2018
Earlier work this paper cites.
Video-to-video synthesis. In Proceedings of the Conference on Neural Information Processing Systems
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. 2018 · 2018
Earlier work this paper cites.
Uni-and-bi-directional video prediction via learning object-centric transformation
Xiongtao Chen and Wenmin Wang. 2019 · 2019
Earlier work this paper cites.
Sparse temporal causal convolution for efficient action modeling. In Proceedings of the ACM International Conference on Multimedia . 592–600
Changmao Cheng, Chi Zhang, Yichen Wei, and Yu-Gang Jiang. 2019 · 2019
Earlier work this paper cites.
Gauge equivariant convolutional networks and the icosahedral CNN. In Proceedings of the International Conference on Machine Learning . 1321–1330
Taco Cohen, Maurice Weiler, Berkay Kicanaoglu, and Max Welling. 2019 · 2019
Cited alongside, same era.
Unsupervised neural quantization for compressed-domain similarity search. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3036–3045
Stanislav Morozov and Artem Babenko. 2019 · 2019
Cited alongside, same era.
Video generation from single semantic label map. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3733–3742
Junting Pan, Chengyu Wang, Xu Jia, Jing Shao, Lu Sheng, Junjie Yan, and Xiaogang Wang. 2019 · 2019
Cited alongside, same era.
Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations
Maziar Raissi, Paris Perdikaris, and George E Karniadakis. 2019 · 2019
Cited alongside, same era.
Revisiting hierarchical approach for persistent long-term video prediction. In Proceedings of the International Conference on Learning Representations
Wonkwang Lee, Whie Jung, Han Zhang, Ting Chen, Jing Yu Koh, Thomas Huang, Hyungsuk Yoon, Honglak Lee, and Seunghoon Hong. 2021 · 2021
Later among the works it cites.
Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators
Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis. 2021 · 2021
Later among the works it cites.
Generating diverse structure for image inpainting with hierarchical VQ-VAE. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10775–10784
Jialun Peng, Dong Liu, Songcen Xu, and Houqiang Li. 2021 · 2021
Later among the works it cites.
Vision transformers for dense prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 12179–12188
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High fidelity video prediction with large stochastic recurrent neural networks. In Proceedings of the Conference on Neural Information Processing Systems
Ruben Villegas, Arkanath Pathak, Harini Kannan, Dumitru Erhan, Quoc V Le, and Honglak Lee. 2019 · 2019
Cited alongside, same era.
Conditional deep surrogate models for stochastic, high-dimensional, and multi-fidelity systems
Yibo Yang and Paris Perdikaris. 2019 · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Stochastic latent residual video prediction. In Proceedings of the International Conference on Machine Learning . 3233–3246
Jean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier, and Patrick Gallinari. 2020 · 2020
Cited alongside, same era.
Modeling the dynamics of PDE systems with physics-constrained deep auto-regressive networks
Nicholas Geneva and Nicholas Zabaras. 2020 · 2020
Cited alongside, same era.
Disentangling physical dynamics from unknown factors for unsupervised video prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11474–11484
Vincent Le Guen and Nicolas Thome. 2020 · 2020
Cited alongside, same era.
Fourier neural operator for parametric partial differential equations. In Proceedings of the International Conference on Learning Representations
Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar. 2020 · 2020
Cited alongside, same era.
Self-attention convlstm for spatiotemporal prediction. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 11531–11538
Zhihui Lin, Maomao Li, Zhuobin Zheng, Yangyang Cheng, and Chun Yuan. 2020 · 2020
Cited alongside, same era.
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. 2021 · 2021
Later among the works it cites.
MotionRNN: A flexible model for video prediction with spacetime-varying motions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15435–15444
Haixu Wu, Zhiyu Yao, Jianmin Wang, and Mingsheng Long. 2021 · 2021
Later among the works it cites.
Videogpt: Video generation using vq-vae and transformers
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. 2021 · 2021
Later among the works it cites.
Strpm: A spatiotemporal residual predictive model for high-resolution video prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13946–13955
Zheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma, and Wen Gao. 2022 · 2022
Later among the works it cites.
Automated discovery of fundamental variables hidden in experimental data
Boyuan Chen, Kuang Huang, Sunand Raghupathi, Ishaan Chandratreya, Qiang Du, and Hod Lipson. 2022 · 2022
Later among the works it cites.
Simvp: Simpler yet better video prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3170–3180
Zhangyang Gao, Cheng Tan, Lirong Wu, and Stan Z Li. 2022 · 2022
Later among the works it cites.
Msdr: Multi-step dependency relation networks for spatial temporal forecasting. In Proceedings of the International ACM SIGKDD Conference on Knowledge Discovery & Data Mining . 1042–1050
Dachuan Liu, Jin Wang, Shuo Shang, and Peng Han. 2022 · 2022
Later among the works it cites.
Predrnn: A recurrent neural network for spatiotemporal predictive learning
Yunbo Wang, Haixu Wu, Jianjin Zhang, Zhifeng Gao, Jianmin Wang, S Yu Philip, and Mingsheng Long. 2022 · 2022
Later among the works it cites.
Optimizing video prediction via video frame interpolation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 17814–17823
Yue Wu, Qiang Wen, and Qifeng Chen. 2022 · 2022
Later among the works it cites.
Active Patterns Perceived for Stochastic Video Prediction. In Proceedings of the ACM International Conference on Multimedia . 5961–5969
Yechao Xu, Zhengxing Sun, Qian Li, Yunhan Sun, and Shoutong Luo. 2022 · 2022
Later among the works it cites.
VPTR: Efficient Transformers for Video Prediction. In International Conference on Pattern Recognition . 3492–3499
Xi Ye and Guillaume-Alexandre Bilodeau. 2022 · 2022
Later among the works it cites.
Efficient VVC Intra Prediction Based on Deep Feature Fusion and Probability Estimation
Tiesong Zhao, Yuhang Huang, Weize Feng, Yiwen Xu, and Sam Kwong. 2022 · 2022
Later among the works it cites.
Robust semantic communications with masked VQ-VAE enabled codebook
Qiyu Hu, Guangyi Zhang, Zhijin Qin, Yunlong Cai, Guanding Yu, and Geoffrey Ye Li. 2023b · 2023
Closest in time.
Revisiting Multi-Codebook Quantization
Xiaosu Zhu, Jingkuan Song, Lianli Gao, Xiaoyan Gu, and Heng Tao Shen. 2023 · 2023
Closest in time.