Fetching the paper…
Reading the bibliography…
Predicting future scene representations is a crucial task for enabling robots to understand and interact with the environment.
The reviewing of object files: Object-specific integration of information
Kahneman, D., Treisman, A., and Gibbs, B. J · 1992
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Wang, Z., Bovik, A. C., Sheikh, H. R., and Simoncelli, E. P · 2004
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Cho, K., van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Automatic differentiation in PyTorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A. and Vinyals, O · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Stochastic video generation with a learned prior
Denton, E. and Fergus, R · 2018
Earlier work this paper cites.
Object perception
Johnson, S. P · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
Monet: Unsupervised scene decomposition and representation
Burgess, C. P., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M., and Lerchner, A · 2019
Earlier work this paper cites.
Imitating latent policies from observation
Edwards, A., Sahni, H., Schroecker, Y., and Isbell, C · 2019
Earlier work this paper cites.
Challenging common assumptions in the unsupervised learning of disentangled representations
Locatello, F., Bauer, S., Lucic, M., Raetsch, G., Gelly, S., Schölkopf, B., and Bachem, O · 2019
Earlier work this paper cites.
Spatial broadcast decoder: A simple architecture for disentangled representations in VAEs, 2019
Watters, N., Matthey, L., Burgess, C. P., and Lerchner, A · 2019
Earlier work this paper cites.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Cabi, S., Colmenarejo, S. G., Novikov, A., Konyushkova, K., Reed, S., Jeong, R., Zolna, K., Aytar, Y., Budden, D., Vecerik, M., et al · 2020
Earlier work this paper cites.
On the binding problem in artificial neural networks
Greff, K., Van Steenkiste, S., and Schmidhuber, J · 2020
Earlier work this paper cites.
Towards practical multi-object manipulation using relational reinforcement learning
Li, R., Jabri, A., Darrell, T., and Agrawal, P · 2020
Earlier work this paper cites.
Object-centric learning with slot attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
Unsupervised object-based transition models for 3D partially observable environments
Creswell, A., Kabra, R., Burgess, C., and Shanahan, M · 2021
Cited alongside, same era.
Playable video generation
Menapace, W., Lathuiliere, S., Tulyakov, S., Siarohin, A., and Ricci, E · 2021
Cited alongside, same era.
Illiterate dall-e learns to compose
Singh, G., Deng, F., and Ahn, S · 2021
Cited alongside, same era.
Parts: Unsupervised segmentation with slots, attention and independence maximization
Zoran, D., Kabra, R., Lerchner, A., and Rezende, D. J · 2021
Cited alongside, same era.
Vim: Variational independent modules for video prediction
Bridging the gap to real-world object-centric learning
Seitzer, M., Horn, M., Zadaianchuk, A., Zietlow, D., Xiao, T., Simon-Gabriel, C.-J., He, T., Zhang, Z., Schölkopf, B., Brox, T., et al · 2023
Later among the works it cites.
ILPO-MP: Mode Priors Prevent Mode Collapse when Imitating Latent Policies from Observations
Struckmeier, O. and Kyrki, V · 2023
Later among the works it cites.
Object-centric video prediction via decoupling of object dynamics and interactions
Villar-Corrales, A., Wahdan, I., and Behnke, S · 2023
Later among the works it cites.
An investigation into pre-training object-centric representations for reinforcement learning
Yoon, J., Wu, Y.-F., Bae, H., and Ahn, S · 2023
Later among the works it cites.
Inverse dynamics pretraining learns good representations for multitask imitation
Brandfonbrener, D., Nachum, O., and Bruna, J · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Assouel, R., Castrejon, L., Courville, A., Ballas, N., and Bengio, Y · 2022
Cited alongside, same era.
Generalization and robustness implications in object-centric learning
Dittadi, A., Papa, S., De Vita, M., Schölkopf, B., Winther, O., and Locatello, F · 2022
Cited alongside, same era.
SAVi++: Towards end-to-end object-centric learning from real-world videos
Elsayed, G. F., Mahendran, A., van Steenkiste, S., Greff, K., Mozer, M. C., and Kipf, T · 2022
Cited alongside, same era.
Conditional Object-Centric Learning from Video
Kipf, T., Elsayed, G. F., Mahendran, A., Stone, A., Sabour, S., Heigold, G., Jonschkowski, R., Dosovitskiy, A., and Greff, K · 2022
Cited alongside, same era.
Playable environments: Video manipulation in space and time
Menapace, W., Lathuilière, S., Siarohin, A., Theobalt, C., Tulyakov, S., Golyanik, V., and Ricci, E · 2022
Cited alongside, same era.
Simple unsupervised object-centric learning for complex and naturalistic videos
Singh, G., Wu, Y.-F., and Ahn, S · 2022
Cited alongside, same era.
Become a proficient player with limited data through watching pure videos
Ye, W., Zhang, Y., Abbeel, P., and Gao, Y · 2022
Cited alongside, same era.
Bruce, J., Dennis, M. D., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., et al · 2024
Later among the works it cites.
DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control
Cui, Z. J., Pan, H., Iyer, A., Haldar, S., and Pinto, L · 2024
Later among the works it cites.
DDLP: Unsupervised Object-centric Video Prediction with Deep Dynamic Latent Particles
Daniel, T. and Tamar, A · 2024
Later among the works it cites.
SlotSSMs: Slot State Space Models
Jiang, J., Deng, F., Singh, G., Lee, M., and Ahn, S · 2024
Later among the works it cites.
Object-centric temporal consistency via conditional autoregressive inductive biases
Meo, C., Nakano, A., Lică, M., Didolkar, A., Suzuki, M., Goyal, A., Zhang, M., Dauwels, J., Matsuo, Y., and Bengio, Y · 2024
Later among the works it cites.
Learning to act without actions
Schmidt, D. and Jiang, M · 2024
Later among the works it cites.
TIV-Diffusion: Towards object-centric movement for text-driven image to video generation
Wang, X., Li, X., Hu, Y., Zhu, H., Hou, C., Lan, C., and Chen, Z · 2024
Later among the works it cites.
Object-centric learning for real-world videos by predicting temporal feature similarities
Zadaianchuk, A., Seitzer, M., and Martius, G · 2024
Later among the works it cites.
Object-centric world model for language-guided manipulation
Jeong, Y., Chun, J., Cha, S., and Kim, T · 2025
Closest in time.
Exploring the effectiveness of object-centric representations in visual question answering: Comparative insights with foundation models
Mamaghan, A. M. K., Papa, S., Johansson, K. H., Bauer, S., and Dittadi, A · 2025
Closest in time.
SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from Pixels
Mosbach, M., Niklas Ewertz, J., Villar-Corrales, A., and Behnke, S · 2025
Closest in time.
Predictive inverse dynamics models are scalable learners for robotic manipulation
Tian, Y., Yang, S., Zeng, J., Wang, P., Lin, D., Dong, H., and Pang, J · 2025
Closest in time.
Object-centric image to video generation with language guidance
Villar-Corrales, A., Plepi, G., and Behnke, S · 2025
Closest in time.
Latent action pretraining from videos
Ye, S., Jang, J., Jeon, B., Joo, S., Yang, J., Peng, B., Mandlekar, A., Tan, R., Chao, Y.-W., Lin, B. Y., et al · 2025
Closest in time.