Fetching the paper…
Reading the bibliography…
Recent years have witnessed remarkable advances in spatiotemporal predictive learning, with methods incorporating auxiliary inputs, complex neural architectures, and sophisticated training strategies.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
C. Schuldt, I. Laptev, and B. Caputo, “Recognizing human actions: a local svm approach,” in Proceedings of the 17th International Conference on Pattern Recognition, 2004. ICPR 2004. , vol. 3. IEEE, 2004, pp. 32–36
2004
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
P. Dollár, C. Wojek, B. Schiele, and P. Perona, “Pedestrian detection: A benchmark,” in 2009 IEEE conference on computer vision and pattern recognition . IEEE, 2009, pp. 304–311
2009
Earlier work this paper cites.
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The kitti dataset,” The International Journal of Robotics Research , vol. 32, no. 11, pp. 1231–1237, 2013
2013
Earlier work this paper cites.
V. Michalski, R. Memisevic, and K. Konda, “Modeling deep temporal dependencies with recurrent grammar cells””,” Advances in neural information processing systems , vol. 27, pp. 1925–1933, 2014
2014
Earlier work this paper cites.
X. Shi, Z. Chen, H. Wang, D.-Y. Yeung, W.-K. Wong, and W.-c. Woo, “Convolutional lstm network: A machine learning approach for precipitation nowcasting,” Advances in Neural Information Processing Systems , vol. 28, 2015
2015
Earlier work this paper cites.
N. Srivastava, E. Mansimov, and R. Salakhudinov, “Unsupervised learning of video representations using lstms,” in International conference on machine learning . PMLR, 2015, pp. 843–852
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 1–9
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
C. Finn, I. Goodfellow, and S. Levine, “Unsupervised learning for physical interaction through video prediction,” Advances in Neural Information Processing Systems , vol. 29, 2016
2016
Earlier work this paper cites.
W. Lotter, G. Kreiman, and D. Cox, “Deep predictive coding networks for video prediction and unsupervised learning,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
X. Jia, B. De Brabandere, T. Tuytelaars, and L. V. Gool, “Dynamic filter networks,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Y. Wang, M. Long, J. Wang, Z. Gao, and P. S. Yu, “Predrnn: Recurrent neural networks for predictive learning using spatiotemporal lstms,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
C. Lu, M. Hirsch, and B. Scholkopf, “Flexible spatio-temporal networks for video prediction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 6523–6531
2017
Earlier work this paper cites.
R. Villegas, J. Yang, Y. Zou, S. Sohn, X. Lin, and H. Lee, “Learning to generate long-term future via hierarchical prediction,” in international conference on machine learning . PMLR, 2017, pp. 3560–3569
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. Villegas, J. Yang, S. Hong, X. Lin, and H. Lee, “Decomposing motion and content for natural video sequence prediction,” in International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
X. Liang, L. Lee, W. Dai, and E. P. Xing, “Dual motion gan for future-flow embedded video prediction,” in proceedings of the IEEE international conference on computer vision , 2017, pp. 1744–1752
2017
Earlier work this paper cites.
E. L. Denton et al. , “Unsupervised learning of disentangled representations from video,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Z. Liu, R. A. Yeh, X. Tang, Y. Liu, and A. Agarwala, “Video frame synthesis using deep voxel flow,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 4463–4471
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
J. Zhang, Y. Zheng, and D. Qi, “Deep spatio-temporal residual networks for citywide crowd flows prediction,” in Thirty-first AAAI conference on artificial intelligence , 2017
2017
Earlier work this paper cites.
N. Kalchbrenner, A. Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. Kavukcuoglu, “Video pixel networks,” in International Conference on Machine Learning . PMLR, 2017, pp. 1771–1779
2017
Earlier work this paper cites.
M. Oliu, J. Selva, and S. Escalera, “Folded recurrent neural networks for future video prediction,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 716–731
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Y. Wang, Z. Gao, M. Long, J. Wang, and S. Y. Philip, “Predrnn++: Towards a resolution of the deep-in-time dilemma in spatiotemporal predictive learning,” in International Conference on Machine Learning . PMLR, 2018, pp. 5123–5132
2018
Earlier work this paper cites.
Y. Wang, L. Jiang, M.-H. Yang, L.-J. Li, M. Long, and L. Fei-Fei, “Eidetic 3d lstm: A model for video prediction and beyond,” in International conference on learning representations , 2018
2018
Earlier work this paper cites.
R. Villegas, D. Erhan, H. Lee et al. , “Hierarchical long-term video prediction without supervision,” in International Conference on Machine Learning . PMLR, 2018, pp. 6038–6046
2018
Earlier work this paper cites.
E. Denton and R. Fergus, “Stochastic video generation with a learned prior,” in International Conference on Machine Learning . PMLR, 2018, pp. 1174–1183
2018
Cited alongside, same era.
Z. Xu, Y. Wang, M. Long, J. Wang, and M. KLiss, “Predcnn: Predictive learning with cascade convolutions.” in IJCAI , 2018, pp. 2940–2947
2018
Cited alongside, same era.
W. Byeon, Q. Wang, R. K. Srivastava, and P. Koumoutsakos, “Contextvp: Fully context-aware video prediction,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 753–769
2018
Cited alongside, same era.
B. Yu, H. Yin, and Z. Zhu, “Spatio-temporal graph convolutional networks: a deep learning framework for traffic forecasting,” in Proceedings of the 27th International Joint Conference on Artificial Intelligence , 2018, pp. 3634–3640
2018
Cited alongside, same era.
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning . PMLR, 2021, pp. 10 347–10 357
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 012–10 022
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Seo, M. Defferrard, P. Vandergheynst, and X. Bresson, “Structured sequence modeling with graph convolutional recurrent networks,” in International conference on neural information processing . Springer, 2018, pp. 362–373
2018
Cited alongside, same era.
Y. Li, R. Yu, C. Shahabi, and Y. Liu, “Diffusion convolutional recurrent neural network: Data-driven traffic forecasting,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
Z. Hao, X. Huang, and S. Belongie, “Controllable video generation with sparse trajectories,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 7854–7863
2018
Cited alongside, same era.
F. A. Reda, G. Liu, K. J. Shih, R. Kirby, J. Barker, D. Tarjan, A. Tao, and B. Catanzaro, “Sdc-net: Video prediction using spatially-displaced convolution,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 718–733
2018
Cited alongside, same era.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 586–595
2018
Cited alongside, same era.
2018
Cited alongside, same era.
B. Jin, Y. Hu, Y. Zeng, Q. Tang, S. Liu, and J. Ye, “Varnet: Exploring variations for unsupervised video prediction,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2018, pp. 5801–5806
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Liu, Z. Dai, D. So, and Q. V. Le, “Pay attention to mlps,” Advances in Neural Information Processing Systems , vol. 34, pp. 9204–9215, 2021
2021
Later among the works it cites.
F. Research, “fvcore,” https://github.com/facebookresearch/fvcore , 2021
2021
Later among the works it cites.
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit et al. , “Mlp-mixer: An all-mlp architecture for vision,” Advances in neural information processing systems , vol. 34, pp. 24 261–24 272, 2021
2021
Later among the works it cites.
W. Wen, W. Ren, Y. Shi, Y. Nie, J. Zhang, and X. Cao, “Video super-resolution via a spatio-temporal alignment network,” IEEE Transactions on Image Processing , vol. 31, pp. 1761–1773, 2022
2022
Closest in time.
Z. Chang, X. Zhang, S. Wang, S. Ma, and W. Gao, “Stam: A spatiotemporal attention based memory for video prediction,” IEEE Transactions on Multimedia , vol. 25, pp. 2354–2367, 2022
2022
Closest in time.
——, “Strpm: A spatiotemporal residual predictive model for high-resolution video prediction,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 13 946–13 955
2022
Closest in time.
Z. Gao, C. Tan, and S. Z. Li, “Simvp: Simpler yet better video prediction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 3170–3180
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
H. Lin, Z. Gao, Y. Xu, L. Wu, L. Li, and S. Z. Li, “Conditional local convolution for spatio-temporal meteorological forecasting,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 7, 2022, pp. 7470–7478
2022
Closest in time.
W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, and S. Yan, “Metaformer is actually what you need for vision,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 10 819–10 829
2022
Closest in time.
A. Trockman and J. Z. Kolter, “Patches are all you need?” arXiv preprint arXiv:2201.09792 , 2022
2022
Closest in time.
Y. Zhang, T. Zhang, C. Wu, and R. Tao, “Multi-scale spatiotemporal feature fusion network for video saliency prediction,” IEEE Transactions on Multimedia , 2023
2023
Closest in time.
Y. Yu, X. Zhao, R. Ni, S. Yang, Y. Zhao, and A. C. Kot, “Augmented multi-scale spatiotemporal inconsistency magnifier for generalized deepfake detection,” IEEE Transactions on Multimedia , vol. 25, pp. 8487–8498, 2023
2023
Closest in time.
J. Liu, Z. Fan, Z. Yang, Y. Su, and X. Yang, “Multi-stage spatio-temporal fusion network for fast and accurate video bit-depth enhancement,” IEEE Transactions on Multimedia , 2023
2023
Closest in time.
P. Li, C. Zhang, and X. Xu, “Fast fourier inception networks for occluded video prediction,” IEEE Transactions on Multimedia , 2023
2023
Closest in time.
L. Yu, Y. Cheng, K. Sohn, J. Lezama, H. Zhang, H. Chang, A. G. Hauptmann, M.-H. Yang, Y. Hao, I. Essa et al. , “Magvit: Masked generative video transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 10 459–10 469
2023
Closest in time.
2023
Closest in time.
S. Tang, C. Li, P. Zhang, and R. Tang, “Swinlstm: Improving spatiotemporal prediction accuracy using swin transformer and lstm,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 13 470–13 479
2023
Closest in time.
X. Hu, Z. Huang, A. Huang, J. Xu, and S. Zhou, “A dynamic multi-scale voxel flow network for video prediction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 6121–6131
2023
Closest in time.
Y. Zhong, L. Liang, I. Zharkov, and U. Neumann, “Mmvp: Motion-matrix-based video prediction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 4273–4283
2023
Closest in time.
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du et al. , “Bridgedata v2: A dataset for robot learning at scale,” in Conference on Robot Learning . PMLR, 2023, pp. 1723–1736
2023
Closest in time.
2024
Closest in time.
Y. Tang, P. Dong, Z. Tang, X. Chu, and J. Liang, “Vmrnn: Integrating vision mamba and lstm for efficient and accurate spatiotemporal forecasting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 5663–5673
2024
Closest in time.
2024
Closest in time.
J. Sun, J. Xie, J.-F. Hu, Z. Lin, J. Lai, W. Zeng, and W.-S. Zheng, “Predicting future instance segmentation with contextual pyramid convlstms,” in Proceedings of the 27th acm international conference on multimedia , 2019, pp. 2043–2051
2051
Closest in time.