Fetching the paper…
Reading the bibliography…
Automatically discovering composable abstractions from raw perceptual data is a long-standing challenge in machine learning.
A parallel computation that assigns canonical object-based frames of reference
Hinton, G. E · 1981
Earlier work this paper cites.
Comparing partitions
Hubert, L. and Arabie, P · 1985
Earlier work this paper cites.
Principal component analysis in application to object orientation
Yi, W. and Marshall, S · 2000
Earlier work this paper cites.
Grandmother cells, symmetry, and invariance: How the term arose and what the facts suggest
Barlow, H · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E · 2010
Earlier work this paper cites.
Transforming auto-encoders
Hinton, G. E., Krizhevsky, A., and Wang, S. D · 2011
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Cho, K., Van Merriënboer, B., Bahdanau, D., and Bengio, Y · 2014
Earlier work this paper cites.
Binding via reconstruction clustering
Greff, K., Srivastava, R. K., and Schmidhuber, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Spatial transformer networks
Jaderberg, M., Simonyan, K., Zisserman, A., and Kavukcuoglu, K · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
Attend, infer, repeat: Fast scene understanding with generative models
Eslami, S. M. A., Heess, N., Weber, T., Tassa, Y., Szepesvari, D., Kavukcuoglu, K., and Hinton, G. E · 2016
Earlier work this paper cites.
Tagger: Deep unsupervised perceptual grouping
Greff, K., Rasmus, A., Berglund, M., Hao, T. H., Valpola, H., and Schmidhuber, J · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Neural expectation maximization
Greff, K., van Steenkiste, S., and Schmidhuber, J · 2017
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C. L., and Girshick, R. B · 2017
Earlier work this paper cites.
SGDR: stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Dynamic routing between capsules
Sabour, S., Frosst, N., and Hinton, G. E · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Matrix capsules with EM routing
Hinton, G. E., Sabour, S., and Frosst, N · 2018
Earlier work this paper cites.
Sequential attend, infer, repeat: Generative modelling of moving objects
Kosiorek, A. R., Kim, H., Teh, Y. W., and Posner, I · 2018
Earlier work this paper cites.
Self-Attention with Relative Position Representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2018
Earlier work this paper cites.
Relational neural expectation maximization: Unsupervised discovery of objects and their interactions
van Steenkiste, S., Chang, M., Greff, K., and Schmidhuber, J · 2018
Earlier work this paper cites.
Group normalization
Wu, Y. and He, K · 2018
Earlier work this paper cites.
Attention Augmented Convolutional Networks
Bello, I., Zoph, B., Le, Q., Vaswani, A., and Shlens, J · 2019
Earlier work this paper cites.
Monet: Unsupervised scene decomposition and representation
Burgess, C. P., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M. M., and Lerchner, A · 2019
Earlier work this paper cites.
Spatially invariant unsupervised object detection with convolutional neural networks
Crawford, E. and Pineau, J · 2019
Cited alongside, same era.
Multi-object representation learning with iterative variational inference
Greff, K., Kaufman, R. L., Kabra, R., Watters, N., Burgess, C., Zoran, D., Matthey, L., Botvinick, M. M., and Lerchner, A · 2019
Cited alongside, same era.
A framework for intelligence and cortical function based on grid cells in the neocortex
Hawkins, J., Lewis, M., Klukas, M., Purdy, S., and Ahmad, S · 2019
Cited alongside, same era.
Multi-object datasets
Kabra, R., Burgess, C., Matthey, L., Kaufman, R. L., Greff, K., Reynolds, M., and Lerchner, A · 2019
Cited alongside, same era.
Stacked capsule autoencoders
Kosiorek, A. R., Sabour, S., Teh, Y. W., and Hinton, G. E · 2019
Cited alongside, same era.
Stand-Alone Self-Attention in Vision Models
Parmar, N., Ramachandran, P., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J · 2019
Unsupervised layered image decomposition into object prototypes
Monnier, T., Vincent, E., Ponce, J., and Aubry, M · 2021
Later among the works it cites.
Vision transformers for dense prediction
Ranftl, R., Bochkovskiy, A., and Koltun, V · 2021
Later among the works it cites.
Sajjadi, M. S. M., Meyer, H., Pot, E., Bergmann, U., Greff, K., Radwan, N., Vora, S., Lucic, M., Duckworth, D., Dosovitskiy, A., Uszkoreit, J., Funkhouser, T. A., and Tagliasacchi, A · 2021
Later among the works it cites.
Illiterate DALL-E learns to compose
Singh, G., Deng, F., and Ahn, S · 2021
Later among the works it cites.
Marionette: Self-supervised sprite learning
Smirnov, D., Gharbi, M., Fisher, M., Guizilini, V., Efros, A. A., and Solomon, J. M · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Spatial broadcast decoder: A simple architecture for learning disentangled representations in vaes
Watters, N., Matthey, L., Burgess, C. P., and Lerchner, A · 2019
Cited alongside, same era.
Knowledge across reference frames: Cognitive maps and image spaces
Bottini, R. and Doeller, C. F · 2020
Cited alongside, same era.
End-to-End Object Detection with Transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Cited alongside, same era.
On the relationship between self-attention and convolutional layers
Cordonnier, J.-B., Loukas, A., and Jaggi, M · 2020
Cited alongside, same era.
GENESIS: generative scene inference and sampling with object-centric latent representations
Engelcke, M., Kosiorek, A. R., Jones, O. P., and Posner, I · 2020
Cited alongside, same era.
On the binding problem in artificial neural networks
Greff, K., Van Steenkiste, S., and Schmidhuber, J · 2020
Cited alongside, same era.
Bottleneck Transformers for Visual Recognition
Srinivas, A., Lin, T.-Y., Parmar, N., Shlens, J., Abbeel, P., and Vaswani, A · 2021
Later among the works it cites.
Decomposing 3d scenes into objects via unsupervised volume segmentation
Stelzner, K., Kersting, K., and Kosiorek, A. R · 2021
Later among the works it cites.
Deformable DETR: Deformable Transformers for End-to-End Object Detection
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J · 2021
Later among the works it cites.
Position Information in Transformers: An Overview
Dufter, P., Schmitt, M., and Schütze, H · 2022
Later among the works it cites.
SAVi++: Towards end-to-end object-centric learning from real-world videos
Elsayed, G. F., Mahendran, A., van Steenkiste, S., Greff, K., Mozer, M. C., and Kipf, T · 2022
Later among the works it cites.
Kubric: A scalable dataset generator
Greff, K., Belletti, F., Beyer, L., Doersch, C., Du, Y., Duckworth, D., Fleet, D. J., Gnanapragasam, D., Golemo, F., Herrmann, C., Kipf, T., Kundu, A., Lagun, D., Laradji, I. H., Liu, H. D., Meyer, H., Miao, Y., Nowrouzezahrai, D., Öztireli, A. C., Pot, E., Radwan, N., Rebain, D., Sabour, S., Sajjadi, M. S. M., Sela, M., Sitzmann, V., Stone, A., Sun, D., Vora, S., Wang, Z., Wu, T., Yi, K. M., Zhong, F., and Tagliasacchi, A · 2022
Later among the works it cites.
Equivariant graph hierarchy-based neural networks
Han, J., Rong, Y., Xu, T., Sun, F., and Huang, W · 2022
Later among the works it cites.
Conditional object-centric learning from video
Kipf, T., Elsayed, G. F., Mahendran, A., Stone, A., Sabour, S., Heigold, G., Jonschkowski, R., Dosovitskiy, A., and Greff, K · 2022
Later among the works it cites.
DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
Liu, S., Li, F., Zhang, H., Yang, X., Qi, X., Su, H., Zhu, J., and Zhang, L · 2022
Later among the works it cites.
Waymo open dataset: Panoramic video panoptic segmentation
Mei, J., Zhu, A. Z., Yan, X., Yan, H., Qiao, S., Chen, L., and Kretzschmar, H · 2022
Later among the works it cites.
Learning symmetric embeddings for equivariant world models
Park, J. Y., Biza, O., Zhao, L., van de Meent, J., and Walters, R · 2022
Later among the works it cites.
Object scene representation transformer
Sajjadi, M. S. M., Duckworth, D., Mahendran, A., van Steenkiste, S., Pavetic, F., Lucic, M., Guibas, L. J., Greff, K., and Kipf, T · 2022
Later among the works it cites.
Unsupervised multi-object segmentation using attention and soft-argmax
Sauvalle, B. and de La Fortelle, A · 2022
Later among the works it cites.
Simple unsupervised object-centric learning for complex and naturalistic videos
Singh, G., Wu, Y., and Ahn, S · 2022
Later among the works it cites.
Anchor DETR: Query Design for Transformer-Based Detector
Wang, Y., Zhang, X., Yang, T., and Sun, J · 2022
Later among the works it cites.
Slotformer: Unsupervised visual dynamics simulation with object-centric models
Wu, Z., Dvornik, N., Greff, K., Kipf, T., and Garg, A · 2022
Later among the works it cites.
Segmenting moving objects via an object-centric layered representation
Xie, J., Xie, W., and Zisserman, A · 2022
Later among the works it cites.
Unsupervised discovery of object radiance fields
Yu, H., Guibas, L. J., and Wu, J · 2022
Later among the works it cites.
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection, 2022
Zhang, H., Li, F., Liu, S., Zhang, L., Su, H., Zhu, J., Ni, L. M., and Shum, H.-Y · 2022
Later among the works it cites.
Slot-VPS: Object-centric representation learning for video panoptic segmentation
Zhou, Y., Zhang, H., Lee, H., Sun, S., Li, P., Zhu, Y., Yoo, B., Qi, X., and Han, J.-J · 2022
Later among the works it cites.
Bridging the gap to real-world object-centric learning
Seitzer, M., Horn, M., Zadaianchuk, A., Zietlow, D., Xiao, T., Simon-Gabriel, C.-J., He, T., Zhang, Z., Schölkopf, B., Brox, T., and Locatello, F · 2023
Closest in time.