Fetching the paper…
Reading the bibliography…
Slot attention has shown remarkable object-centric representation learning performance in computer vision tasks without requiring any supervision.
Mental models: Towards a cognitive science of language, inference, and consciousness
Johnson-Laird, P. N · 1983
Earlier work this paper cites.
Comparing partitions
Hubert, L. and Arabie, P · 1985
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Vision as bayesian inference: analysis by synthesis?
Yuille, A. and Kersten, D · 2006
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Geiger, A., Lenz, P., and Urtasun, R · 2012
Earlier work this paper cites.
Simulation as an engine of physical scene understanding
Battaglia, P. W., Hamrick, J. B., and Tenenbaum, J. B · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B · 2016
Earlier work this paper cites.
Attend, infer, repeat: Fast scene understanding with generative models
Eslami, S., Heess, N., Weber, T., Tassa, Y., Szepesvari, D., Hinton, G. E., et al · 2016
Earlier work this paper cites.
Neural expectation maximization
Greff, K., Van Steenkiste, S., and Schmidhuber, J · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Earlier work this paper cites.
A simple neural network module for relational reasoning
Santoro, A., Raposo, D., Barrett, D. G., Malinowski, M., Pascanu, R., Battaglia, P., and Lillicrap, T · 2017
Earlier work this paper cites.
Isolating sources of disentanglement in variational autoencoders
Chen, R. T. Q., Li, X., Grosse, R., and Duvenaud, D · 2018
Earlier work this paper cites.
Deep object-centric representations for generalizable robot learning
Devin, C., Abbeel, P., Darrell, T., and Levine, S · 2018
Earlier work this paper cites.
Shapestacks: Learning vision-based physical intuition for generalised object stacking
Groth, O., Fuchs, F. B., Posner, I., and Vedaldi, A · 2018
Earlier work this paper cites.
Ha, D. and Schmidhuber, J · 2018
Cited alongside, same era.
Sequential attend, infer, repeat: Generative modelling of moving objects
Kosiorek, A., Kim, H., Teh, Y. W., and Posner, I · 2018
Cited alongside, same era.
Rezende, D. J. and Viola, F · 2018
Cited alongside, same era.
Monet: Unsupervised scene decomposition and representation
Burgess, C. P., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M., and Lerchner, A · 2019
Cited alongside, same era.
Spatially invariant unsupervised object detection with convolutional neural networks
Crawford, E. and Pineau, J · 2019
Cited alongside, same era.
Object-centric learning with slot attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Later among the works it cites.
Blockgan: Learning 3d object-aware scene representations from unlabelled images
Nguyen-Phuoc, T. H., Richardt, C., Mai, L., Yang, Y., and Mitra, N · 2020
Later among the works it cites.
Investigating object compositionality in generative adversarial networks
Van Steenkiste, S., Kurach, K., Schmidhuber, J., and Gelly, S · 2020
Later among the works it cites.
Generative scene graph networks
Deng, F., Zhi, Z., Lee, D., and Ahn, S · 2021
Later among the works it cites.
Efficient iterative amortized inference for learning symmetric and disentangled multi-object representations
Emami, P., He, P., Ranka, S., and Rangarajan, A · 2021
Later among the works it cites.
Genesis-v2: Inferring unordered object representations without iterative refinement
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Engelcke, M., Kosiorek, A. R., Jones, O. P., and Posner, I · 2019
Cited alongside, same era.
Cyclical annealing schedule: A simple approach to mitigating kl vanishing
Fu, H., Li, C., Liu, X., Gao, J., Celikyilmaz, A., and Carin, L · 2019
Cited alongside, same era.
Multi-object representation learning with iterative variational inference
Greff, K., Kaufman, R. L., Kabra, R., Watters, N., Burgess, C., Zoran, D., Matthey, L., Botvinick, M., and Lerchner, A · 2019
Cited alongside, same era.
Scalor: Generative world models with scalable object representations
Jiang, J., Janghorbani, S., De Melo, G., and Ahn, S · 2019
Cited alongside, same era.
Multi-object datasets
Kabra, R., Burgess, C., Matthey, L., Kaufman, R. L., Greff, K., Reynolds, M., and Lerchner, A · 2019
Cited alongside, same era.
Spatial broadcast decoder: A simple architecture for learning disentangled representations in vaes
Watters, N., Matthey, L., Burgess, C. P., and Lerchner, A · 2019
Cited alongside, same era.
Relate: Physically plausible multi-object scene synthesis using structured latent spaces
Ehrhardt, S., Groth, O., Monszpart, A., Engelcke, M., Posner, I., Mitra, N., and Vedaldi, A · 2020
Cited alongside, same era.
Engelcke, M., Parker Jones, O., and Posner, I · 2021
Later among the works it cites.
Conditional object-centric learning from video
Kipf, T., Elsayed, G. F., Mahendran, A., Stone, A., Sabour, S., Heigold, G., Jonschkowski, R., Dosovitskiy, A., and Greff, K · 2021
Later among the works it cites.
Giraffe: Representing scenes as compositional generative neural feature fields
Niemeyer, M. and Geiger, A · 2021
Later among the works it cites.
Toward causal representation learning
Schölkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y · 2021
Later among the works it cites.
Illiterate dall-e learns to compose
Singh, G., Deng, F., and Ahn, S · 2021
Later among the works it cites.
Greedy hierarchical variational autoencoders for large-scale video prediction
Wu, B., Nair, S., Martin-Martin, R., Fei-Fei, L., and Finn, C · 2021
Later among the works it cites.
Savi++: Towards end-to-end object-centric learning from real-world videos
Elsayed, G. F., Mahendran, A., van Steenkiste, S., Greff, K., Mozer, M. C., and Kipf, T · 2022
Later among the works it cites.
Slot order matters for compositional scene understanding
Emami, P., He, P., Ranka, S., and Rangarajan, A · 2022
Later among the works it cites.
Compositional multi-object reinforcement learning with linear relation networks
Mambelli, D., Träuble, F., Bauer, S., Schölkopf, B., and Locatello, F · 2022
Later among the works it cites.
Bridging the gap to real-world object-centric learning
Seitzer, M., Horn, M., Zadaianchuk, A., Zietlow, D., Xiao, T., Simon-Gabriel, C.-J., He, T., Zhang, Z., Schölkopf, B., Brox, T., et al · 2022
Later among the works it cites.
Neural systematic binder, 2022
Singh, G., Kim, Y., and Ahn, S · 2022
Later among the works it cites.
Jiang, J., Deng, F., Singh, G., and Ahn, S · 2023
Closest in time.