Fetching the paper…
Reading the bibliography…
Unsupervised multi-object scene decomposition is a fast-emerging problem in representation learning.
Marr, D.: A theory for cerebral neocortex. Proceedings of the Royal Society of London
1970
Earlier work this paper cites.
Rand, W.M.: Objective criteria for the evaluation of clustering methods. Journal of the American Statistical association
1971
Earlier work this paper cites.
Hubert, L., Arabie, P.: Comparing partitions. Journal of classification
1985
Earlier work this paper cites.
Hinton, G., Sejnowski, T.: Learning and relearning in boltzmann machines. Parallel Distributed Processing: Explorations in the Microstructure of Cognition
1986
Earlier work this paper cites.
Weiss, Y., Adelson, E.: A unified mixture framework for mo- tion segmentation: Incorporating spatial coherence and estimating the number of models. In: CVPR. pp. 321–326 (1996)
1996
Earlier work this paper cites.
Hochreiter, S., Schmidhuber, J.: Long short-term memory. Neural computation
1997
Earlier work this paper cites.
Shi, J., Malik, J.: Normalized cuts and image segmentation. IEEE Transactions on pattern analysis and machine intelligence
2000
Earlier work this paper cites.
Jepson, A., Fleet, D., Black, M.: A layered motion representation with occlusion and compact spatial support. In: ECCV. pp. 692–706 (2002)
2002
Earlier work this paper cites.
Felzenszwalb, P.F., Huttenlocher, D.P.: Efficient graph-based image segmentation. International journal of computer vision
2004
Earlier work this paper cites.
Graves, A., Fernández, S., Schmidhuber, J.: Bidirectional lstm networks for improved phoneme classification and recognition. In: Duch, W., Kacprzyk, J., Oja, E., Zadrożny, S. (eds.) Artificial Neural Networks: Formal Models and Their Applications – ICANN 2005. pp. 799–804. Springer Berlin Heidelberg, Berlin, Heidelberg (2005)
2005
Earlier work this paper cites.
Graves, A., Fernández, S., Schmidhuber, J.: Multi-dimensional recurrent neural networks. In: International conference on artificial neural networks. pp. 549–558. Springer (2007)
2007
Earlier work this paper cites.
Vincent, P., Larochelle, H., Bengio, Y., Manzagol, P.A.: Extracting and composing robust features with denoising autoencoders. In: ICML (2008)
2008
Earlier work this paper cites.
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., Bengio, Y.: Learning phrase representations using rnn encoder-decoder for statistical machine translation. In: EMNLP (2014)
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Kingma, D.P., Welling, M.: Auto-encoding variational bayes. In: ICLR (2014)
2014
Earlier work this paper cites.
Chung, J., Kastner, K., Dinh, L., Goel, K., Courville, A.C., Bengio, Y.: A recurrent latent variable model for sequential data. In: NIPS. pp. 2980–2988 (2015)
2015
Earlier work this paper cites.
Girshick, R.: Fast r-cnn. In: ICCV. pp. 1440–1448 (2015)
2015
Earlier work this paper cites.
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M.W., Pfau, D., Schaul, T., de Freitas, N.: Learning to learn by gradient descent by gradient descent. In: NIPS (2016)
2016
Cited alongside, same era.
Eslami, S.A., Heess, N., Weber, T., Tassa, Y., Szepesvari, D., Hinton, G.E., et al.: Attend, infer, repeat: Fast scene understanding with generative models. In: NIPS. pp. 3225–3233 (2016)
2016
Cited alongside, same era.
Greff, K., Rasmus, A., Berglund, M., Hao, T., Valpola, H., Schmidhuber, J.: Tagger: Deep unsupervised perceptual grouping. In: NIPS. pp. 4484–4492 (2016)
2016
Cited alongside, same era.
Noroozi, M., Favaro, P.: Unsupervised learning of visual representations by solving jigsaw puzzles. In: ECCV (2016)
2016
Cited alongside, same era.
Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., Efros, A.A.: Context encoders: Feature learning by inpainting. In: CVPR (2016)
Marino, J., Yue, Y., Mandt, S.: Iterative amortized inference. In: ICML (2018)
2018
Later among the works it cites.
Redmon, J., Farhadi, A.: Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767 (2018)
2018
Later among the works it cites.
Van Steenkiste, S., Chang, M., Greff, K., Schmidhuber, J.: Relational neural expectation maximization: Unsupervised discovery of objects and their interactions. In: ICLR (2018)
2018
Later among the works it cites.
2019
Later among the works it cites.
Crawford, E., Pineau, J.: Exploiting spatial invariance for scalable unsupervised object tracking. In: AAAI (2019)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Vondrick, C., Pirsiavash, H., Torralba, A.: Anticipating visual representations with unlabeled videos. In: CVPR (2016)
2016
Cited alongside, same era.
Zhang, R., Isola, P., Efros, A.A.: Colorful image colorization. In: ECCV (2016)
2016
Cited alongside, same era.
Chung, J., Ahn, S., Bengio, Y.: Hierarchical multiscale recurrent neural networks. In: ICLR (2017)
2017
Cited alongside, same era.
Cremer, C., Li, X., Duvenaud, D.: Inference suboptimality in variational autoencoders. In: NIPS Workshop on Advances in Approximate Bayesian Inference (2017)
2017
Cited alongside, same era.
Greff, K., Van Steenkiste, S., Schmidhuber, J.: Neural expectation maximization. In: NIPS. pp. 6691–6701 (2017)
2017
Cited alongside, same era.
He, K., Gkioxari, G., Dollár, P., Girshick, R.: Mask r-cnn. In: ICCV. pp. 2961–2969 (2017)
2017
Cited alongside, same era.
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., Lerchner, A.: beta-vae: Learning basic visual concepts with a constrained variational framework. In: ICLR (2017)
2017
Cited alongside, same era.
2019
Later among the works it cites.
Crawford, E., Pineau, J.: Spatially invariant unsupervised object detection with convolutional neural networks. In: AAAI. pp. 3412–3420 (2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Lu, J., Batra, D., Parikh, D., Lee, S.: Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In: NeurIPS (2019)
2019
Later among the works it cites.
Sun, C., Myers, A., Vondrick, C., Murphy, K., Schmidt, C.: Videobert: A joint model for video and language representation learning. In: ICCV (2019)
2019
Later among the works it cites.
Tan, H., Bansal, M.: Lxmert: Learning cross-modality encoder representations from transformers. In: Conference on Empirical Methods in Natural Language Processing (2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Casanova, A., Pinheiro, P.O., Rostamzadeh, N., Pal, C.J.: Reinforced active learning for image segmentation. In: ICLR (2020)
2020
Closest in time.
Engelcke, M., Kosiorek, A.R., Jones, O.P., Posner, I.: Genesis: Generative scene inference and sampling with object-centric latent representations. In: ICLR (2020)
2020
Closest in time.
Jiang, J., Janghorbani, S., De Melo, G., Ahn, S.: Scalor: Generative world models with scalable object representations. In: ICLR (2020)
2020
Closest in time.
Kossen, J., Stelzner, K., Hussing, M., Voelcker, C., Kersting, K.: Structured object-aware physics prediction for video modeling and planning. In: ICLR (2020)
2020
Closest in time.
Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., Tenenbaum, J.B.: Clevrer: Collision events for video representation and reasoning. In: ICLR (2020)
2020
Closest in time.