Fetching the paper…
Reading the bibliography…
The ability to carve the world into useful abstractions in order to reason about time and space is a crucial component of intelligence.
Monet: Unsupervised scene decomposition and representation
Christopher P. Burgess, Loïc Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matthew Botvinick, and Alexander Lerchner · 1901
Earlier work this paper cites.
Reducing BERT pre-training time from 3 days to 76 minutes
Yang You, Jing Li, Jonathan Hseu, Xiaodan Song, James Demmel, and Cho-Jui Hsieh · 1904
Earlier work this paper cites.
An introduction to variational autoencoders
Diederik P. Kingma and Max Welling · 1906
Earlier work this paper cites.
CATER: A diagnostic dataset for compositional actions and temporal reasoning
Rohit Girdhar and Deva Ramanan · 1910
Earlier work this paper cites.
On Information and Sufficiency
S. Kullback and R. A. Leibler · 1951
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Object-centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2006
Earlier work this paper cites.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge J. Belongie, and Yin Cui · 2008
Earlier work this paper cites.
Attention over learned object embeddings enables complex visual reasoning, 2020
David Ding, Felix Hill, Adam Santoro, Malcolm Reynolds, and Matt Botvinick · 2012
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent, 2016
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W. Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando de Freitas · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning, 2016
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2016
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
João Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Cited alongside, same era.
The "something something" video database for learning and evaluating visual common sense, 2017
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzyńska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, and Roland Memisevic · 2017
Cited alongside, same era.
A simple neural network module for relational reasoning, 2017
Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Hierarchical conditional relation networks for video question answering, 2020
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran · 2020
Later among the works it cites.
Clevrer: Collision events for video representation and reasoning, 2020
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes, 2020
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Later among the works it cites.
Vivit: A video vision transformer, 2021
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
João Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman · 2018
Cited alongside, same era.
Relational inductive bias for physical construction in humans and machines
Jessica B. Hamrick, Kelsey R. Allen, Victor Bapst, Tina Zhu, Kevin R. McKee, Joshua B. Tenenbaum, and Peter W. Battaglia · 2018
Cited alongside, same era.
Compositional attention networks for machine reasoning, 2018
Drew A. Hudson and Christopher D. Manning · 2018
Cited alongside, same era.
Joseph Marino, Yisong Yue, and Stephan Mandt · 2018
Cited alongside, same era.
Relational recurrent neural networks, 2018
Adam Santoro, Ryan Faulkner, David Raposo, Jack Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, and Timothy Lillicrap · 2018
Cited alongside, same era.
Tvqa: Localized, compositional video question answering, 2019
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L. Berg · 2019
Cited alongside, same era.
Stabilizing transformers for reinforcement learning, 2019
Emilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant M. Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, Matthew M. Botvinick, Nicolas Heess, and Raia Hadsell · 2019
Cited alongside, same era.
Spatial broadcast decoder: A simple architecture for learning disentangled representations in vaes, 2019
Nicholas Watters, Loic Matthey, Christopher P. Burgess, and Alexander Lerchner · 2019
Cited alongside, same era.
Grounding physical concepts of objects and events through dynamic visual reasoning, 2021
Zhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong, Joshua B. Tenenbaum, and Chuang Gan · 2021
Later among the works it cites.
Tutorial on variational autoencoders, 2021
Carl Doersch · 2021
Later among the works it cites.
A large-scale study on unsupervised spatiotemporal representation learning, 2021
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He · 2021
Later among the works it cites.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and João Carreira · 2021
Later among the works it cites.
Conditional object-centric learning from video
Thomas Kipf, Gamaleldin F. Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, and Klaus Greff · 2021
Later among the works it cites.
Movinets: Mobile video networks for efficient video recognition, 2021
Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, and Boqing Gong · 2021
Later among the works it cites.
Compositional processing emerges in neural networks solving math problems, 2021
Jacob Russin, Roland Fernandez, Hamid Palangi, Eric Rosen, Nebojsa Jojic, Paul Smolensky, and Jianfeng Gao · 2021
Later among the works it cites.
Long-short temporal contrastive learning of video transformers
Jue Wang, Gedas Bertasius, Du Tran, and Lorenzo Torresani · 2021
Later among the works it cites.
Unsupervised learning of temporal abstractions with slot-based transformers, 2022
Anand Gopalakrishnan, Kazuki Irie, Jürgen Schmidhuber, and Sjoerd van Steenkiste · 2022
Closest in time.