Fetching the paper…
Reading the bibliography…
We present Language-mediated, Object-centric Representation Learning (LORL), a paradigm for learning disentangled, object-centric scene representations from vision and language.
Monet: Unsupervised scene decomposition and representation
Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner. 2019 · 1901
Earlier work this paper cites.
Objective criteria for the evaluation of clustering methods
William M Rand. 1971 · 1971
Earlier work this paper cites.
Comparing partitions
Lawrence Hubert and Phipps Arabie. 1985 · 1985
Earlier work this paper cites.
Object individuation and object identity in infancy: The role of spatiotemporal information, object property information, and language
Fei Xu. 1999 · 1999
Earlier work this paper cites.
How Children Learn the Meanings of Words
Paul Bloom. 2002 · 2002
Earlier work this paper cites.
Object files and schemata: Factorizing declarative and procedural knowledge in dynamical systems
Anirudh Goyal, Alex Lamb, Phanideep Gampa, Philippe Beaudoin, Sergey Levine, Charles Blundell, Yoshua Bengio, and Michael Mozer. 2020 · 2006
Earlier work this paper cites.
Sortal concepts, object individuation, and language
Fei Xu. 2007 · 2007
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
From perceptual to language-mediated categorization
Gert Westermann and Denis Mareschal. 2014 · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Attend, infer, repeat: Fast scene understanding with generative models
SM Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, Koray Kavukcuoglu, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Tagger: Deep unsupervised perceptual grouping
Klaus Greff, Antti Rasmus, Mathias Berglund, Tele Hao, Harri Valpola, and Jürgen Schmidhuber. 2016 · 2016
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Joint image-text representation by gaussian visual-semantic embedding
Zhou Ren, Hailin Jin, Zhe Lin, Chen Fang, and Alan Yuille. 2016 · 2016
Cited alongside, same era.
Neural expectation maximization
Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Neural scene de-rendering
Jiajun Wu, Joshua B Tenenbaum, and Pushmeet Kohli. 2017 · 2017
Cited alongside, same era.
Obj2Text: Generating visually descriptive language from object layouts
Xuwang Yin and Vicente Ordonez. 2017 · 2017
Cited alongside, same era.
Vse++: Improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018 · 2018
Cited alongside, same era.
Learning by abstraction: The neural state machine
Drew Hudson and Christopher D Manning. 2019 · 2019
Later among the works it cites.
Clevr-ref+: Diagnosing visual reasoning with referring expressions
Runtao Liu, Chenxi Liu, Yutong Bai, and Alan L Yuille. 2019 · 2019
Later among the works it cites.
The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision
Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B. Tenenbaum, and Jiajun Wu. 2019 · 2019
Later among the works it cites.
Partnet: A large-scale benchmark for fine-grained and hierarchical part-level 3d object understanding
Kaichun Mo, Shilin Zhu, Angel X Chang, Li Yi, Subarna Tripathi, Leonidas J Guibas, and Hao Su. 2019 · 2019
Later among the works it cites.
Faster attend-infer-repeat with tractable probabilistic models
Karl Stelzner, Robert Peharz, and Kristian Kersting. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unsupervised video object segmentation for deep reinforcement learning
Vikash Goel, Jameson Weng, and Pascal Poupart. 2018 · 2018
Cited alongside, same era.
Sequential attend, infer, repeat: Generative modelling of moving objects
Adam R. Kosiorek, Hyunjik Kim, Ingmar Posner, and Yee Whye Teh. 2018 · 2018
Cited alongside, same era.
Object counts! bringing explicit detections back into image captioning
Josiah Wang, Pranava Swaroop Madhyastha, and Lucia Specia. 2018 · 2018
Cited alongside, same era.
Visual curiosity: Learning to ask questions to learn visual recognition
Jianwei Yang, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh. 2018 · 2018
Cited alongside, same era.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum. 2018 · 2018
Cited alongside, same era.
Shapeglot: Learning language for shape differentiation
Panos Achlioptas, Judy Fan, Robert Hawkins, Noah Goodman, and Leonidas J Guibas. 2019 · 2019
Cited alongside, same era.
Daniel M Bear, Chaofei Fan, Damian Mrowca, Yunzhu Li, Seth Alter, Aran Nayebi, Jeremy Schwartz, Li Fei-Fei, Jiajun Wu, Joshua B Tenenbaum, et al. 2020 · 2020
Closest in time.
Genesis: Generative scene inference and sampling with object-centric latent representations
Martin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, and Ingmar Posner. 2020 · 2020
Closest in time.
Contrastive learning of structured world models
Thomas Kipf, Elise van der Pol, and Max Welling. 2020 · 2020
Closest in time.
Space: Unsupervised object-oriented scene representation via spatial attention and decomposition
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun, Gautam Singh, Fei Deng, Jindong Jiang, and Sungjin Ahn. 2020 · 2020
Closest in time.
Object-centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. 2020 · 2020
Closest in time.
Learning representations from audio-visual spatial alignment
Pedro Morgado, Yi Li, and Nuno Nvasconcelos. 2020 · 2020
Closest in time.
Shaping visual representations with language for few-shot classification
Jesse Mu, Percy Liang, and Noah Goodman. 2020 · 2020
Closest in time.
Shop-vrb: A visual reasoning benchmark for object perception
M. Nazarczuk and K. Mikolajczyk. 2020 · 2020
Closest in time.
Embodied language grounding with 3d visual feature representations
Mihir Prabhudesai, H. F. Tung, Syed Ashar Javed, Maximilian Sieb, Adam W. Harley, and K. Fragkiadaki. 2020 · 2020
Closest in time.
Entity abstraction in visual model-based reinforcement learning
Rishi Veerapaneni, John D Co-Reyes, Michael Chang, Michael Janner, Chelsea Finn, Jiajun Wu, Joshua Tenenbaum, and Sergey Levine. 2020 · 2020
Closest in time.