Fetching the paper…
Reading the bibliography…
Spatiotemporal predictive learning aims to generate future frames by learning from historical frames.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Recognizing human actions: a local svm approach
Christian Schuldt, Ivan Laptev, and Barbara Caputo · 2004
Earlier work this paper cites.
Pedestrian detection: A benchmark
Piotr Dollár, Christian Wojek, Bernt Schiele, and Pietro Perona · 2009
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
Video (language) modeling: a baseline for generative models of natural videos
Arthur Szlam Marc’Aurelio Ranzato, Joan Bruna, Michaël Mathieu, Ronan Collobert, and Sumit Chopra · 2014
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Dynamic filter networks
Xu Jia, Bert De Brabandere, Tinne Tuytelaars, and Luc V Gool · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2016
Earlier work this paper cites.
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A Efros · 2016
Earlier work this paper cites.
Unsupervised learning of disentangled representations from video
Emily L Denton et al · 2017
Earlier work this paper cites.
Density estimation using real nvp
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio · 2017
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017
Earlier work this paper cites.
Video pixel networks
Nal Kalchbrenner, Aäron Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Dual motion gan for future-flow embedded video prediction
Xiaodan Liang, Lisa Lee, Wei Dai, and Eric P Xing · 2017
Earlier work this paper cites.
Video frame synthesis using deep voxel flow
Ziwei Liu, Raymond A Yeh, Xiaoou Tang, Yiming Liu, and Aseem Agarwala · 2017
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
William Lotter, Gabriel Kreiman, and David Cox · 2017
Earlier work this paper cites.
Deep learning for precipitation nowcasting: A benchmark and a new model
Xingjian Shi, Zhihan Gao, Leonard Lausen, Hao Wang, Dit-Yan Yeung, Wai-kin Wong, and Wang-chun Woo · 2017
Earlier work this paper cites.
Decomposing motion and content for natural video sequence prediction
Ruben Villegas, Jimei Yang, Seunghoon Hong, Xunyu Lin, and Honglak Lee · 2017
Earlier work this paper cites.
Predrnn: Recurrent neural networks for predictive learning using spatiotemporal lstms
Yunbo Wang, Mingsheng Long, Jianmin Wang, Zhifeng Gao, and Philip S Yu · 2017
Earlier work this paper cites.
Deep spatio-temporal residual networks for citywide crowd flows prediction
Junbo Zhang, Yu Zheng, and Dekang Qi · 2017
Earlier work this paper cites.
Learning spatiotemporal features using 3dcnn and convolutional lstm for gesture recognition
Liang Zhang, Guangming Zhu, Peiyi Shen, Juan Song, Syed Afaq Shah, and Mohammed Bennamoun · 2017
Earlier work this paper cites.
Stochastic variational video prediction
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H Campbell, and Sergey Levine · 2018
Earlier work this paper cites.
Contextvp: Fully context-aware video prediction
Wonmin Byeon, Qin Wang, Rupesh Kumar Srivastava, and Petros Koumoutsakos · 2018
Earlier work this paper cites.
Stochastic video generation with a learned prior
Emily Denton and Rob Fergus · 2018
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations
Spyros Gidaris, Praveer Singh, and Nikos Komodakis · 2018
Earlier work this paper cites.
Controllable video generation with sparse trajectories
Zekun Hao, Xun Huang, and Serge Belongie · 2018
Earlier work this paper cites.
Learning to decompose and disentangle representations for video prediction
Jun-Ting Hsieh, Bingbin Liu, De-An Huang, Li F Fei-Fei, and Juan Carlos Niebles · 2018
Earlier work this paper cites.
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun · 2018
Earlier work this paper cites.
Varnet: Exploring variations for unsupervised video prediction
Beibei Jin, Yu Hu, Yiming Zeng, Qiankun Tang, Shice Liu, and Jing Ye · 2018
Cited alongside, same era.
Stochastic adversarial video prediction
Alex X Lee, Richard Zhang, Frederik Ebert, Pieter Abbeel, Chelsea Finn, and Sergey Levine · 2018
Cited alongside, same era.
Mutual suppression network for video prediction using disentangled features
Jungbeom Lee, Jangho Lee, Sungmin Lee, and Sungroh Yoon · 2018
Cited alongside, same era.
Folded recurrent neural networks for future video prediction
Marc Oliu, Javier Selva, and Sergio Escalera · 2018
Cited alongside, same era.
Sdc-net: Video prediction using spatially-displaced convolution
Fitsum A Reda, Guilin Liu, Kevin J Shih, Robert Kirby, Jon Barker, David Tarjan, Andrew Tao, and Bryan Catanzaro · 2018
Cited alongside, same era.
Variational autoencoders and nonlinear ica: A unifying framework
Ilyes Khemakhem, Diederik Kingma, Ricardo Monti, and Aapo Hyvarinen · 2020
Later among the works it cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Later among the works it cites.
Trajectorycnn: a new spatio-temporal feature learning network for human motion prediction
Xiaoli Liu, Jianqin Yin, Jin Liu, Pengxiang Ding, Jun Liu, and Huaping Liu · 2020
Later among the works it cites.
A review on deep learning techniques for video prediction
Sergiu Oprea, Pablo Martinez-Gonzalez, Alberto Garcia-Garcia, John Alejandro Castro-Vargas, Sergio Orts-Escolano, Jose Garcia-Rodriguez, and Antonis Argyros · 2020
Later among the works it cites.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical long-term video prediction without supervision
Ruben Villegas, Dumitru Erhan, Honglak Lee, et al · 2018
Cited alongside, same era.
Rgb-d-based human motion recognition with deep learning: A survey
Pichao Wang, Wanqing Li, Philip Ogunbona, Jun Wan, and Sergio Escalera · 2018
Cited alongside, same era.
Predrnn++: Towards a resolution of the deep-in-time dilemma in spatiotemporal predictive learning
Yunbo Wang, Zhifeng Gao, Mingsheng Long, Jianmin Wang, and S Yu Philip · 2018
Cited alongside, same era.
Eidetic 3d lstm: A model for video prediction and beyond
Yunbo Wang, Lu Jiang, Ming-Hsuan Yang, Li-Jia Li, Mingsheng Long, and Li Fei-Fei · 2018
Cited alongside, same era.
Predcnn: Predictive learning with cascade convolutions
Ziru Xu, Yunbo Wang, Mingsheng Long, Jianmin Wang, and M KLiss · 2018
Cited alongside, same era.
Improved conditional vrnns for video prediction
Lluis Castrejon, Nicolas Ballas, and Aaron Courville · 2019
Cited alongside, same era.
Gstnet: Global spatial-temporal network for traffic flow prediction
Shen Fang, Qi Zhang, Gaofeng Meng, Shiming Xiang, and Chunhong Pan · 2019
Cited alongside, same era.
Osamu Shouno · 2020
Later among the works it cites.
Convolutional tensor-train lstm for spatio-temporal learning
Jiahao Su, Wonmin Byeon, Jean Kossaifi, Furong Huang, Jan Kautz, and Anima Anandkumar · 2020
Later among the works it cites.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Tongzhou Wang and Phillip Isola · 2020
Later among the works it cites.
Fitvid: Overfitting in pixel-level video prediction
Mohammad Babaeizadeh, Mohammad Taghi Saffar, Suraj Nair, Sergey Levine, Chelsea Finn, and Dumitru Erhan · 2021
Later among the works it cites.
Graph and temporal convolutional networks for 3d multi-person pose estimation in monocular videos
Yu Cheng, Bo Wang, Bo Yang, and Robby T Tan · 2021
Later among the works it cites.
How to represent part-whole hierarchies in a neural network
Geoffrey Hinton · 2021
Later among the works it cites.
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Later among the works it cites.
Video prediction recalling long-term motion context via memory alignment learning
Sangmin Lee, Hak Gu Kim, Dae Hwi Choi, Hyung-Il Kim, and Yong Man Ro · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui · 2021
Later among the works it cites.
Co-learning: Learning from noisy labels with self-supervision
Cheng Tan, Jun Xia, Lirong Wu, and Stan Z Li · 2021
Later among the works it cites.
Predrnn: A recurrent neural network for spatiotemporal predictive learning
Yunbo Wang, Haixu Wu, Jianjin Zhang, Zhifeng Gao, Jianmin Wang, Philip S Yu, and Mingsheng Long · 2021
Later among the works it cites.
Motionrnn: A flexible model for video prediction with spacetime-varying motions
Haixu Wu, Zhiyu Yao, Jianmin Wang, and Mingsheng Long · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny · 2021
Later among the works it cites.
Enhancing neural sign language translation by highlighting the facial expression information
Jiangbin Zheng, Yidong Chen, Chong Wu, Xiaodong Shi, and Suhail Muhammad Kamal · 2021
Later among the works it cites.
A survey on generative diffusion model
Hanqun Cao, Cheng Tan, Zhangyang Gao, Guangyong Chen, Pheng-Ann Heng, and Stan Z Li · 2022
Closest in time.
Scaling up your kernels to 31x31: Revisiting large kernel design in cnns
Xiaohan Ding, Xiangyu Zhang, Yizhuang Zhou, Jungong Han, Guiguang Ding, and Jian Sun · 2022
Closest in time.
Alphadesign: A graph protein design method and benchmark on alphafolddb
Zhangyang Gao, Cheng Tan, Stan Li, et al · 2022
Closest in time.
Pifold: Toward effective and efficient protein inverse folding
Zhangyang Gao, Cheng Tan, and Stan Z Li · 2022
Closest in time.
Simvp: Simpler yet better video prediction
Zhangyang Gao, Cheng Tan, Lirong Wu, and Stan Z. Li · 2022
Closest in time.
Meng-Hao Guo, Cheng-Ze Lu, Zheng-Ning Liu, Ming-Ming Cheng, and Shi-Min Hu · 2022
Closest in time.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Closest in time.
Efficient multi-order gated aggregation network
Siyuan Li, Zedong Wang, Zicheng Liu, Cheng Tan, Haitao Lin, Di Wu, Zhiyuan Chen, Jiangbin Zheng, and Stan Z Li · 2022
Closest in time.
Decoupled mixup for data-efficient learning
Zicheng Liu, Siyuan Li, Ge Wang, Cheng Tan, Lirong Wu, and Stan Z Li · 2022
Closest in time.
Automix: Unveiling the power of mixup for stronger classifiers
Zicheng Liu, Siyuan Li, Di Wu, Zihan Liu, Zhiyuan Chen, Lirong Wu, and Stan Z Li · 2022
Closest in time.
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Closest in time.
Rfold: Towards simple yet effective rna secondary structure prediction
Cheng Tan, Zhangyang Gao, and Stan Z Li · 2022
Closest in time.
Target-aware molecular graph generation
Cheng Tan, Zhangyang Gao, and Stan Z Li · 2022
Closest in time.
Hyperspherical consistency regularization
Cheng Tan, Zhangyang Gao, Lirong Wu, Siyuan Li, and Stan Z Li · 2022
Closest in time.
Flowformer: Linearizing transformers with conservation flows
Haixu Wu, Jialong Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long · 2022
Closest in time.
Using context-to-vector with graph retrofitting to improve word embeddings
Jiangbin Zheng, Yile Wang, Ge Wang, Jun Xia, Yufei Huang, Guojiang Zhao, Yue Zhang, and Stan Li · 2022
Closest in time.