Fetching the paper…
Reading the bibliography…
Spatio-temporal predictive learning is a learning paradigm that enables models to learn spatial and temporal patterns by predicting future frames from given past frames in an unsupervised manner.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Recognizing human actions: a local svm approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
Pedestrian detection: A benchmark
P. Dollár, C. Wojek, B. Schiele, and P. Perona · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun · 2013
Earlier work this paper cites.
Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu · 2013
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
X. Shi, Z. Chen, H. Wang, D.-Y. Yeung, W.-K. Wong, and W.-c. Woo · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhudinov · 2015
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Video frame synthesis using deep voxel flow
Z. Liu, R. A. Yeh, X. Tang, Y. Liu, and A. Agarwala · 2017
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
W. Lotter, G. Kreiman, and D. Cox · 2017
Earlier work this paper cites.
Predrnn: Recurrent neural networks for predictive learning using spatiotemporal lstms
Y. Wang, M. Long, J. Wang, Z. Gao, and P. S. Yu · 2017
Earlier work this paper cites.
Deep spatio-temporal residual networks for citywide crowd flows prediction
J. Zhang, Y. Zheng, and D. Qi · 2017
Earlier work this paper cites.
Learning spatiotemporal features using 3dcnn and convolutional lstm for gesture recognition
L. Zhang, G. Zhu, P. Shen, J. Song, S. Afaq Shah, and M. Bennamoun · 2017
Earlier work this paper cites.
S. Aigner and M. Körner · 2018
Earlier work this paper cites.
MMCV: OpenMMLab computer vision foundation
M. Contributors · 2018
Earlier work this paper cites.
Learning to decompose and disentangle representations for video prediction
J.-T. Hsieh, B. Liu, D.-A. Huang, L. F. Fei-Fei, and J. C. Niebles · 2018
Earlier work this paper cites.
Folded recurrent neural networks for future video prediction
M. Oliu, J. Selva, and S. Escalera · 2018
Earlier work this paper cites.
Hierarchical long-term video prediction without supervision
R. Villegas, D. Erhan, H. Lee, et al · 2018
Earlier work this paper cites.
Rgb-d-based human motion recognition with deep learning: A survey
P. Wang, W. Li, P. Ogunbona, J. Wan, and S. Escalera · 2018
Cited alongside, same era.
Predrnn++: Towards a resolution of the deep-in-time dilemma in spatiotemporal predictive learning
Y. Wang, Z. Gao, M. Long, J. Wang, and S. Y. Philip · 2018
Cited alongside, same era.
Eidetic 3d lstm: A model for video prediction and beyond
Y. Wang, L. Jiang, M.-H. Yang, L.-J. Li, M. Long, and L. Fei-Fei · 2018
Cited alongside, same era.
Predcnn: Predictive learning with cascade convolutions
Z. Xu, Y. Wang, M. Long, J. Wang, and M. KLiss · 2018
Cited alongside, same era.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Cited alongside, same era.
Improved conditional vrnns for video prediction
L. Castrejon, N. Ballas, and A. Courville · 2019
Predrnn: A recurrent neural network for spatiotemporal predictive learning
Y. Wang, H. Wu, J. Zhang, Z. Gao, J. Wang, P. S. Yu, and M. Long · 2021
Later among the works it cites.
Self-supervised learning on graphs: Contrastive, generative, or predictive
L. Wu, H. Lin, C. Tan, Z. Gao, and S. Z. Li · 2021
Later among the works it cites.
Simvp: Simpler yet better video prediction
Z. Gao, C. Tan, and S. Z. Li · 2022
Later among the works it cites.
M.-H. Guo, C.-Z. Lu, Z.-N. Liu, M.-M. Cheng, and S.-M. Hu · 2022
Later among the works it cites.
Uniformer: Unifying convolution and self-attention for visual recognition
K. Li, Y. Wang, J. Zhang, P. Gao, G. Song, Y. Liu, H. Li, and Y. Qiao · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gstnet: Global spatial-temporal network for traffic flow prediction
S. Fang, Q. Zhang, G. Meng, S. Xiang, and C. Pan · 2019
Cited alongside, same era.
Deep learning and process understanding for data-driven earth system science
M. Reichstein, G. Camps-Valls, B. Stevens, M. Jung, J. Denzler, N. Carvalhais, et al · 2019
Cited alongside, same era.
Memory in memory: A predictive neural network for learning higher-order non-stationarity from spatiotemporal dynamics
Y. Wang, J. Zhang, H. Zhu, M. Long, J. Wang, and P. S. Yu · 2019
Cited alongside, same era.
Efficient and information-preserving future frame prediction and beyond
W. Yu, Y. Lu, S. Easterbrook, and S. Fidler · 2019
Cited alongside, same era.
Disentangling physical dynamics from unknown factors for unsupervised video prediction
V. L. Guen and N. Thome · 2020
Cited alongside, same era.
Video representation learning by recognizing temporal transformations
S. Jenni, G. Meishvili, and P. Favaro · 2020
Cited alongside, same era.
Efficient multi-order gated aggregation network
S. Li, Z. Wang, Z. Liu, C. Tan, H. Lin, D. Wu, Z. Chen, J. Zheng, and S. Z. Li · 2022
Later among the works it cites.
Openmixup: Open mixup toolbox and benchmark for visual representation learning
S. Li, Z. Wang, Z. Liu, D. Wu, C. Tan, and S. Z. Li · 2022
Later among the works it cites.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie · 2022
Later among the works it cites.
Hornet: Efficient high-order spatial interactions with recursive gated convolutions
Y. Rao, W. Zhao, Y. Tang, J. Zhou, S. N. Lim, and J. Lu · 2022
Later among the works it cites.
Simvp: Towards simple yet powerful spatiotemporal predictive learning
C. Tan, Z. Gao, S. Li, and S. Z. Li · 2022
Later among the works it cites.
A. Trockman and J. Z. Kolter · 2022
Later among the works it cites.
Usb: A unified semi-supervised learning benchmark for classification
Y. Wang, H. Chen, Y. Fan, W. Sun, R. Tao, W. Hou, R. Wang, L. Yang, Z. Zhou, L.-Z. Guo, et al · 2022
Later among the works it cites.
Knowledge distillation improves graph structure augmentation for graph neural networks
L. Wu, H. Lin, Y. Huang, and S. Z. Li · 2022
Later among the works it cites.
Graphmixup: Improving class-imbalanced node classification by reinforcement mixup and self-supervised context prediction
L. Wu, J. Xia, Z. Gao, H. Lin, C. Tan, and S. Z. Li · 2022
Later among the works it cites.
Metaformer is actually what you need for vision
W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, and S. Yan · 2022
Later among the works it cites.
A dynamic multi-scale voxel flow network for video prediction
X. Hu, Z. Huang, A. Huang, J. Xu, and S. Zhou · 2023
Closest in time.
Implicit stacked autoregressive model for video prediction
M. Seo, H. Lee, D. Kim, and J. Seo · 2023
Closest in time.
Temporal attention unit: Towards efficient spatiotemporal predictive learning
C. Tan, Z. Gao, L. Wu, Y. Xu, J. Xia, S. Li, and S. Z. Li · 2023
Closest in time.
Pastnet: Introducing physical inductive biases for spatio-temporal video prediction
H. Wu, W. Xion, F. Xu, X. Luo, C. Chen, X.-S. Hua, and H. Wang · 2023
Closest in time.