Fetching the paper…
Reading the bibliography…
To help agents reason about scenes in terms of their building blocks, we wish to extract the compositional structure of any given scene (in particular, the configuration and characteristics of objects comprising the scene).
Exploiting spatial invariance for scalable unsupervised object tracking
Eric Crawford and Joelle Pineau · 1911
Earlier work this paper cites.
Simultaneous localization and mapping: part i
H. Durrant-Whyte and T. Bailey · 2006
Earlier work this paper cites.
Simultaneous localization and mapping (slam): part ii
T. Bailey and H. Durrant-Whyte · 2006
Earlier work this paper cites.
View invariance for human action recognition
Vasu Parameswaran and Rama Chellappa · 2006
Earlier work this paper cites.
The slam problem: A survey
Josep Aulinas, Yvan Petillot, Joaquim Salvi, and Xavier Lladó · 2008
Earlier work this paper cites.
End-to-End video instance segmentation with transformers
Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia · 2011
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
A benchmark for the evaluation of rgb-d slam systems
J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers · 2012
Earlier work this paper cites.
MaX-DeepLab: End-to-End panoptic segmentation with mask transformers
Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen · 2012
Earlier work this paper cites.
Scene understanding in the era of deep learning, Jun 2015
Jitendra Malik · 2015
Earlier work this paper cites.
Past, present, and future of simultaneous localization and mapping: Towards the robust-perception age
C. Cadena, L. Carlone, H. Carrillo, Y. Latif, D. Scaramuzza, J. Neira, I. Reid, and J.J. Leonard · 2016
Earlier work this paper cites.
Attend, infer, repeat: fast scene understanding with generative models
S M Ali Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, David Szepesvari, Koray Kavukcuoglu, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
The euroc micro aerial vehicle datasets
Michael Burri, Janosch Nikolic, Pascal Gohl, Thomas Schneider, Joern Rehder, Sammy Omari, Markus W Achtelik, and Roland Siegwart · 2016
Earlier work this paper cites.
Chapter 20 - Scene Understanding Using Deep Learning , pages 373–382
Farzad Husain, Babette Dellen, and Carme Torras · 2017
Earlier work this paper cites.
Unsupervised learning of disentangled representations from video
Emily L Denton and Vighnesh Birodkar · 2017
Earlier work this paper cites.
Decomposing motion and content for natural video sequence prediction
Ruben Villegas, Jimei Yang, Seunghoon Hong, Xunyu Lin, and Honglak Lee · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Neural scene representation and rendering
S M Ali Eslami, Danilo Jimenez Rezende, Frederic Besse, Fabio Viola, Ari S Morcos, Marta Garnelo, Avraham Ruderman, Andrei A Rusu, Ivo Danihelka, Karol Gregor, David P Reichert, Lars Buesing, Theophane Weber, Oriol Vinyals, Dan Rosenbaum, Neil Rabinowitz, Helen King, Chloe Hillier, Matt Botvinick, Daan Wierstra, Koray Kavukcuoglu, and Demis Hassabis · 2018
Earlier work this paper cites.
Sequential attend, infer, repeat: Generative modelling of moving objects
Adam Kosiorek, Hyunjik Kim, Yee Whye Teh, and Ingmar Posner · 2018
Earlier work this paper cites.
Danilo Jimenez Rezende and Fabio Viola · 2018
Earlier work this paper cites.
Hierarchically learned view-invariant representations for cross-view action recognition
Yang Liu, Zhaoyang Lu, Jing Li, and Tao Yang · 2018
Earlier work this paper cites.
Stochastic video generation with a learned prior
Emily Denton and Rob Fergus · 2018
Earlier work this paper cites.
Disentangled sequential autoencoder
Li Yingzhen and Stephan Mandt · 2018
Earlier work this paper cites.
Datasets used to train generative query networks (gqns) in the ‘neural scene representation and rendering’ paper
Fabio Viola, Louise Deason, and Marcel Büsching · 2018
Cited alongside, same era.
Monet: Unsupervised scene decomposition and representation
Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner · 2019
Cited alongside, same era.
Multi-object representation learning with iterative variational inference
Klaus Greff, Raphaël Lopez Kaufman, Rishabh Kabra, Nick Watters, Christopher Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, and Alexander Lerchner · 2019
Cited alongside, same era.
Scene representation networks: Continuous 3D-Structure-Aware neural scene representations
Vincent Sitzmann, Michael Zollhoefer, and Gordon Wetzstein · 2019
Cited alongside, same era.
Mid-fusion: Octree-based object-level multi-instance dynamic slam
Binbin Xu, Wenbin Li, Dimos Tzoumanikas, Michael Bloesch, Andrew Davison, and Stefan Leutenegger · 2019
GRF: Learning a general radiance field for 3D scene representation and rendering
Alex Trevithick and Bo Yang · 2020
Later among the works it cites.
pixelNeRF: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa · 2020
Later among the works it cites.
GIRAFFE: Representing scenes as compositional generative neural feature fields
Michael Niemeyer and Andreas Geiger · 2020
Later among the works it cites.
Deepslam: A robust monocular slam system with unsupervised deep learning
Ruihao Li, Sen Wang, and Dongbing Gu · 2020
Later among the works it cites.
Unsupervised learning-based depth estimation-aided visual slam approach
Mingyang Geng, Suning Shang, Bo Ding, Huaimin Wang, and Pengfei Zhang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
SCALOR: Generative world models with scalable object representations
Jindong Jiang, Sepehr Janghorbani, Gerard de Melo, and Sungjin Ahn · 2019
Cited alongside, same era.
An introduction to variational autoencoders
Diederik P Kingma and Max Welling · 2019
Cited alongside, same era.
Multi-object datasets
Rishabh Kabra, Chris Burgess, Loic Matthey, Raphael Lopez Kaufman, Klaus Greff, Malcolm Reynolds, and Alexander Lerchner · 2019
Cited alongside, same era.
Spatial broadcast decoder: A simple architecture for learning disentangled representations in vaes
Nicholas Watters, Loic Matthey, Christopher P Burgess, and Alexander Lerchner · 2019
Cited alongside, same era.
Genesis: Generative scene inference and sampling with object-centric latent representations
Martin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, and Ingmar Posner · 2020
Cited alongside, same era.
Space: Unsupervised object-oriented scene representation via spatial attention and decomposition
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun, Gautam Singh, Fei Deng, Jindong Jiang, and Sungjin Ahn · 2020
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Cited alongside, same era.
Later among the works it cites.
CATER: A diagnostic dataset for compositional actions & temporal reasoning
Rohit Girdhar and Deva Ramanan · 2020
Later among the works it cites.
Imitating interactive intelligence
Josh Abramson, Arun Ahuja, Arthur Brussee, Federico Carnevale, Mary Cassin, Stephen Clark, Andrew Dudzik, Petko Georgiev, Aurelia Guy, Tim Harley, et al · 2020
Later among the works it cites.
View-invariant, occlusion-robust probabilistic embedding for human pose
Ting Liu, Jennifer J Sun, Long Zhao, Jiaping Zhao, Liangzhe Yuan, Yuxiao Wang, Liang-Chieh Chen, Florian Schroff, and Hartwig Adam · 2020
Later among the works it cites.
View-invariant probabilistic embedding for human pose
Jennifer J Sun, Jiaping Zhao, Liang-Chieh Chen, Florian Schroff, Hartwig Adam, and Ting Liu · 2020
Later among the works it cites.
Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction
Yana Hasson, Bugra Tekin, Federica Bogo, Ivan Laptev, Marc Pollefeys, and Cordelia Schmid · 2020
Later among the works it cites.
S3VAE: Self-supervised sequential VAE for representation disentanglement and data generation
Yizhe Zhu, Martin Renqiang Min, Asim Kadav, and Hans Peter Graf · 2020
Later among the works it cites.
End-to-End object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
Rethinking semantic segmentation from a Sequence-to-Sequence perspective with transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip H S Torr, and Li Zhang · 2020
Later among the works it cites.
Grounded language learning fast and slow
Felix Hill, Olivier Tieleman, Tamara von Glehn, Nathaniel Wong, Hamza Merzic, and Stephen Clark · 2020
Later among the works it cites.
Unsupervised object-based transition models for 3d partially observable environments
Antonia Creswell, Rishabh Kabra, Chris Burgess, and Murray Shanahan · 2021
Closest in time.
BARF: Bundle-Adjusting neural radiance fields
Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Simon Lucey · 2021
Closest in time.
NeRF–: Neural radiance fields without known camera parameters
Zirui Wang, Shangzhe Wu, Weidi Xie, Min Chen, and Victor Adrian Prisacariu · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Closest in time.
Is Space-Time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann · 2021
Closest in time.
ViViT: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Closest in time.