Fetching the paper…
Reading the bibliography…
Predicting future dynamics is crucial for applications like autonomous driving and robotics, where understanding the environment is key.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Cross-stitch networks for multi-task learning
Ishan Misra, Abhinav Shrivastava, Abhinav Gupta, and Martial Hebert · 2016
Earlier work this paper cites.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille · 2017
Earlier work this paper cites.
Learning to act by predicting the future
Alexey Dosovitskiy and Vladlen Koltun · 2017
Earlier work this paper cites.
Video scene parsing with predictive feature learning
Xiaojie Jin, Xin Li, Huaxin Xiao, Xiaohui Shen, Zhe Lin, Jimei Yang, Yunpeng Chen, Jian Dong, Luoqi Liu, Zequn Jie, et al · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Earlier work this paper cites.
Fully-adaptive feature sharing in multi-task networks with applications in person attribute classification
Yongxi Lu, Abhishek Kumar, Shuangfei Zhai, Yu Cheng, Tara Javidi, and Rogerio Feris · 2017
Earlier work this paper cites.
Predicting deeper into the future of semantic segmentation
Pauline Luc, Natalia Neverova, Camille Couprie, Jakob Verbeek, and Yann LeCun · 2017
Earlier work this paper cites.
Pyramid scene parsing network
Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia · 2017
Earlier work this paper cites.
Stochastic variational video prediction
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H. Campbell, and Sergey Levine · 2018
Earlier work this paper cites.
Adversarial training for multi-context joint entity and relation extraction
Giannis Bekoulis, Johannes Deleu, Thomas Demeester, and Chris Develder · 2018
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2018
Earlier work this paper cites.
Stochastic adversarial video prediction
Alex X Lee, Richard Zhang, Frederik Ebert, Pieter Abbeel, Chelsea Finn, and Sergey Levine · 2018
Earlier work this paper cites.
Predicting future instance segmentation by forecasting convolutional features
Pauline Luc, Camille Couprie, Yann Lecun, and Jakob Verbeek · 2018
Earlier work this paper cites.
Future semantic segmentation with convolutional lstm
Seyed Shahabeddin Nabavi, Mrigank Rochan, and Yang Wang · 2018
Earlier work this paper cites.
Multi-task learning as multi-objective optimization
Ozan Sener and Vladlen Koltun · 2018
Earlier work this paper cites.
Future semantic segmentation using 3d structure
Suhani Vora, Reza Mahjourian, Soeren Pirk, and Anelia Angelova · 2018
Earlier work this paper cites.
Eidetic 3d lstm: A model for video prediction and beyond
Yunbo Wang, Lu Jiang, Ming-Hsuan Yang, Li-Jia Li, Mingsheng Long, and Li Fei-Fei · 2018
Earlier work this paper cites.
Structure preserving video prediction
Jingwei Xu, Bingbing Ni, Zefan Li, Shuo Cheng, and Xiaokang Yang · 2018
Earlier work this paper cites.
Bayesian prediction of future street scenes using synthetic likelihoods
Apratim Bhattacharyya, Mario Fritz, and Bernt Schiele · 2019
Earlier work this paper cites.
Stochastic filter groups for multi-task CNNs: Learning specialist and generalist convolution kernels
Felix JS Bragman, Ryutaro Tanno, Sebastien Ourselin, Daniel C. Alexander, and Jorge Cardoso · 2019
Earlier work this paper cites.
Improved conditional vrnns for video prediction
Lluis Castrejon, Nicolas Ballas, and Aaron Courville · 2019
Earlier work this paper cites.
Multi-timescale context encoding for scene parsing prediction
Xin Chen and Yahong Han · 2019
Earlier work this paper cites.
Attentive single-tasking of multiple tasks
Kevis-Kokitsi Maninis, Ilija Radosavovic, and Iasonas Kokkinos · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron van den Oord, and Oriol Vinyals · 2019
Earlier work this paper cites.
Latent multi-task architecture learning
Sebastian Ruder, Joachim Bingel, Isabelle Augenstein, and Anders Søgaard · 2019
Earlier work this paper cites.
Single level feature-to-feature forecasting with deformable convolutions
Josip Šarić, Marin Oršić, Tonći Antunović, Sacha Vražić, and Siniša Šegvić · 2019
Earlier work this paper cites.
Predicting future instance segmentation with contextual pyramid convlstms
Jiangxin Sun, Jiafeng Xie, Jian-Fang Hu, Zihang Lin, Jianhuang Lai, Wenjun Zeng, and Wei-Shi Zheng · 2019
Earlier work this paper cites.
Recurrent flow-guided semantic forecasting
Adam Terwilliger, Garrick Brazil, and Xiaoming Liu · 2019
Earlier work this paper cites.
Fixing the train-test resolution discrepancy
Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Herve Jegou · 2019
Earlier work this paper cites.
Detectron2
Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick · 2019
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
Segmenting the future
Hsu-kuang Chiu, Ehsan Adeli, and Juan Carlos Niebles · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Bootstrap your own latent a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko · 2020
Cited alongside, same era.
Learning to branch for multi-task learning
Pengsheng Guo, Chen-Yu Lee, and Daniel Ulbricht · 2020
Cited alongside, same era.
A pytorch reproduction of masked generative image transformer
Victor Besnier and Mickael Chen · 2023
Later among the works it cites.
Dynamic neural network for multi-task learning searching across diverse network topologies
Wonhyeok Choi and Sunghoon Im · 2023
Later among the works it cites.
Eva: Exploring the limits of masked visual representation learning at scale
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao · 2023
Later among the works it cites.
OmniMAE: Single Model Masked Pretraining on Images and Videos
Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Later among the works it cites.
Maskvit: Masked visual pre-training for video prediction
Agrim Gupta, Stephen Tian, Yunzhi Zhang, Jiajun Wu, Roberto Martín-Martín, and Li Fei-Fei · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anthony Hu, Fergal Cotter, Nikhil Mohan, Corina Gurau, and Alex Kendall · 2020
Cited alongside, same era.
Big transfer (bit): General visual representation learning
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby · 2020
Cited alongside, same era.
Warp to the future: Joint forecasting of features and feature motion
Josip Saric, Marin Orsic, Tonci Antunovic, Sacha Vrazic, and Sinisa Segvic · 2020
Cited alongside, same era.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Cited alongside, same era.
Panoptic Segmentation Forecasting
Colin Graber, Grace Tsai, Michael Firman, Gabriel Brostow, and Alexander Schwing · 2021
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang · 2023
Later among the works it cites.
Gaia-1: A generative world model for autonomous driving
Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, and Gianluca Corrado · 2023
Later among the works it cites.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Later among the works it cites.
AdaMTL: Adaptive Input-dependent Inference for Efficient Multi-Task Learning
Marina Neseem, Ahmed Agiza, and Sherief Reda · 2023
Later among the works it cites.
Hiera: a hierarchical vision transformer without the bells-and-whistles
Chaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei, Haoqi Fan, Po-Yao Huang, Vaibhav Aggarwal, Arkabandhu Chowdhury, Omid Poursaeed, Judy Hoffman, Jitendra Malik, Yanghao Li, and Christoph Feichtenhofer · 2023
Later among the works it cites.
Eva-clip: Improved training techniques for clip at scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao · 2023
Later among the works it cites.
Videomae v2: Scaling video masked autoencoders with dual masking
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao · 2023
Later among the works it cites.
Magvit: Masked generative video transformer
Lijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama, Han Zhang, Huiwen Chang, Alexander G Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, et al · 2023
Later among the works it cites.
Anticipative feature fusion transformer for multi-modal action anticipation
Zeyun Zhong, David Schneider, Michael Voit, Rainer Stiefelhagen, and Jürgen Beyerer · 2023
Later among the works it cites.
4m-21: An any-to-any vision model for tens of tasks and modalities
Roman Bachmann, Oğuzhan Fatih Kar, David Mizrahi, Ali Garjani, Mingfei Gao, David Griffiths, Jiaming Hu, Afshin Dehghan, and Amir Zamir · 2024
Closest in time.
Revisiting feature prediction for learning visual representations from video
Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mido Assran, and Nicolas Ballas · 2024
Closest in time.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al · 2024
Closest in time.
Vision transformers need registers
Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski · 2024
Closest in time.
Eva-02: A visual representation for neon genesis
Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao · 2024
Closest in time.
Vista: A generalizable driving world model with high fidelity and versatile controllability
Shenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta, Yihang Qiu, Andreas Geiger, Jun Zhang, and Hongyang Li · 2024
Closest in time.
MOCA: Self-supervised representation learning by predicting masked online codebook assignments
Spyros Gidaris, Andrei Bursuc, Oriane Siméoni, Antonín Vobecký, Nikos Komodakis, Matthieu Cord, and Patrick Perez · 2024
Closest in time.
Spot: Self-training with patch-order permutation for object-centric learning with autoregressive transformers
Ioannis Kakogeorgiou, Spyros Gidaris, Konstantinos Karantzalos, and Nikos Komodakis · 2024
Closest in time.
Videopoet: A large language model for zero-shot video generation
Dan Kondratyuk, Lijun Yu, Xiuye Gu, Jose Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Xuan Yang, Hartwig Adam, Ming-Hsuan Yang, Irfan Essa, Huisheng Wang, David A Ross, Bryan Seybold, and Lu Jiang · 2024
Closest in time.
Autoregressive image generation without vector quantization
Tianhong Li, Yonglong Tian, He Li, Mingyang Deng, and Kaiming He · 2024
Closest in time.
DINOv2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2024
Closest in time.
Givt: Generative infinite-vocabulary transformers
Michael Tschannen, Cian Eastwood, and Fabian Mentzer · 2024
Closest in time.
Emu3: Next-token prediction is all you need
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, et al · 2024
Closest in time.
Language model beats diffusion - tokenizer is key to visual generation
Lijun Yu, Jose Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A Ross, and Lu Jiang · 2024
Closest in time.
Genad: Generative end-to-end autonomous driving
Wenzhao Zheng, Ruiqi Song, Xianda Guo, Chenming Zhang, and Long Chen · 2024
Closest in time.
Lotus: Diffusion-based visual foundation model for high-quality dense prediction
Jing He, Haodong LI, Wei Yin, Yixun Liang, Leheng Li, Kaiqiang Zhou, Hongbo Zhang, Bingbing Liu, and Ying-Cong Chen · 2025
Closest in time.
Radiov2.5: Improved baselines for agglomerative vision foundation models
Greg Heinrich, Mike Ranzinger, Hongxu Yin, Yao Lu, Jan Kautz, Andrew Tao, Bryan Catanzaro, and Pavlo Molchanov · 2025
Closest in time.
Advancing semantic future prediction through multimodal visual sequence transformers
Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, and Nikos Komodakis · 2025
Closest in time.
Dip: Unsupervised dense in-context post-training of visual representations
Sophia Sirko-Galouchenko, Spyros Gidaris, Antonin Vobecky, Andrei Bursuc, and Nicolas Thome · 2025
Closest in time.
Franca: Nested matryoshka clustering for scalable visual representation learning
Shashanka Venkataramanan, Valentinos Pariza, Mohammadreza Salehi, Lukas Knobel, Spyros Gidaris, Elias Ramzi, Andrei Bursuc, and Yuki M Asano · 2025
Closest in time.
DINO-WM: World models on pre-trained visual features enable zero-shot planning
Gaoyue Zhou, Hengkai Pan, Yann LeCun, and Lerrel Pinto · 2025
Closest in time.