Fetching the paper…
Reading the bibliography…
Learning to model how the world changes as time elapses has proven a challenging problem for the computer vision community.
Learning classification with unlabeled data
Virginia R de Sa · 1994
Earlier work this paper cites.
Recursive estimation of generative models of video
Nemanja Petrovic, Aleksandar Ivanovic, and Nebojsa Jojic · 2006
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
A data-driven approach for event prediction
Jenny Yuen and Antonio Torralba · 2010
Earlier work this paper cites.
Activity forecasting
Kris M Kitani, Brian D Ziebart, James Andrew Bagnell, and Martial Hebert · 2012
Earlier work this paper cites.
Action-reaction: Forecasting the dynamics of human interaction
De-An Huang and Kris M Kitani · 2014
Earlier work this paper cites.
Joint summarization of large-scale collections of web images and videos for storyline reconstruction
Gunhee Kim, Leonid Sigal, and Eric P Xing · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
A hierarchical representation for future action prediction
Tian Lan, Tsung-Chuan Chen, and Silvio Savarese · 2014
Earlier work this paper cites.
Seeing the arrow of time
Lyndsey C Pickup, Zheng Pan, Donglai Wei, YiChang Shih, Changshui Zhang, Andrew Zisserman, Bernhard Scholkopf, and William T Freeman · 2014
Earlier work this paper cites.
Video (language) modeling: a baseline for generative models of natural videos
MarcAurelio Ranzato, Arthur Szlam, Joan Bruna, Michael Mathieu, Ronan Collobert, and Sumit Chopra · 2014
Earlier work this paper cites.
Patch to the future: Unsupervised visual prediction
Jacob Walker, Abhinav Gupta, and Martial Hebert · 2014
Earlier work this paper cites.
Instructional videos for unsupervised harvesting and learning of action examples
Shoou-I Yu, Lu Jiang, and Alexander Hauptmann · 2014
Earlier work this paper cites.
Learning image representations tied to ego-motion
Dinesh Jayaraman and Kristen Grauman · 2015
Earlier work this paper cites.
What’s cookin’? interpreting cooking videos using text, speech and vision
Jonathan Malmaud, Jonathan Huang, Vivek Rathod, Nick Johnston, Andrew Rabinovich, and Kevin Murphy · 2015
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2015
Earlier work this paper cites.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using LSTMs
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Dense optical flow prediction from a static image
Jacob Walker, Abhinav Gupta, and Martial Hebert · 2015
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
End-to-end neural sentence ordering using pointer network
Jingjing Gong, Xinchi Chen, Xipeng Qiu, and Xuanjing Huang · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Sentence ordering and coherence modeling using recurrent neural networks
Lajanugen Logeswaran, Honglak Lee, and Dragomir Radev · 2016
Cited alongside, same era.
Deep predictive coding networks for video prediction and unsupervised learning
William Lotter, Gabriel Kreiman, and David Cox · 2016
Cited alongside, same era.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Cited alongside, same era.
Anticipating visual representations from unlabeled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
Stochastic video generation with a learned prior
Emily Denton and Rob Fergus · 2018
Later among the works it cites.
Learning to decompose and disentangle representations for video prediction
Jun-Ting Hsieh, Bingbin Liu, De-An Huang, Li Fei-Fei, and Juan Carlos Niebles · 2018
Later among the works it cites.
Stochastic adversarial video prediction
Alex X Lee, Richard Zhang, Frederik Ebert, Pieter Abbeel, Chelsea Finn, and Sergey Levine · 2018
Later among the works it cites.
Learning and using the arrow of time
Donglai Wei, Joseph J Lim, Andrew Zisserman, and William T Freeman · 2018
Later among the works it cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
MSR-VTT: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Cited alongside, same era.
Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks
Tianfan Xue, Jiajun Wu, Katherine Bouman, and Bill Freeman · 2016
Cited alongside, same era.
Learning dense correspondence via 3D-guided cycle consistency
Tinghui Zhou, Philipp Krähenbühl, Mathieu Aubry, Qixing Huang, and Alexei A. Efros · 2016
Cited alongside, same era.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Cited alongside, same era.
Stochastic variational video prediction
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H Campbell, and Sergey Levine · 2017
Cited alongside, same era.
Self-supervised visual planning with temporal skip connections
Frederik Ebert, Chelsea Finn, Alex X Lee, and Sergey Levine · 2017
Cited alongside, same era.
Temporal cycle-consistency learning
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zisserman · 2019
Later among the works it cites.
What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention
Antonino Furnari and Giovanni Maria Farinella · 2019
Later among the works it cites.
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman · 2019
Later among the works it cites.
Time-agnostic prediction: Predicting predictable video frames
Dinesh Jayaraman, Frederik Ebert, Alexei A Efros, and Sergey Levine · 2019
Later among the works it cites.
Canonical surface mapping via geometric cycle consistency
Nilesh Kulkarni, Abhinav Gupta, and Shubham Tulsiani · 2019
Later among the works it cites.
Videoflow: A conditional flow-based model for stochastic video generation
Manoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn, Sergey Levine, Laurent Dinh, and Durk Kingma · 2019
Later among the works it cites.
Leveraging the present to anticipate the future in videos
Antoine Miech, Ivan Laptev, Josef Sivic, Heng Wang, Lorenzo Torresani, and Du Tran · 2019
Later among the works it cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Later among the works it cites.
Learning video representations using contrastive bidirectional transformer
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
Stochastic prediction of multi-agent interactions from partial observations
Chen Sun, Per Karlsson, Jiajun Wu, Joshua B Tenenbaum, and Kevin Murphy · 2019
Later among the works it cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
Learning correspondence from the cycle-consistency of time
Xiaolong Wang, Allan Jabri, and Alexei A Efros · 2019
Later among the works it cites.
Cross-task weakly supervised learning from instructional videos
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic · 2019
Later among the works it cites.
Memory-augmented dense predictive coding for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman · 2020
Later among the works it cites.
Space-time correspondence as a contrastive random walk
Allan Jabri, Andrew Owens, and Alexei A Efros · 2020
Later among the works it cites.
SLM: Learning a discourse language representation with sentence unshuffling
Haejun Lee, Drew A Hudson, Kangwook Lee, and Christopher D Manning · 2020
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Later among the works it cites.
Learning to anticipate egocentric actions by imagination
Yu Wu, Linchao Zhu, Xiaohan Wang, Yi Yang, and Fei Wu · 2020
Later among the works it cites.