Fetching the paper…
Reading the bibliography…
Physical reasoning is a crucial aspect in the development of general AI systems, given that human learning starts with interacting with the physical world before progressing to more complex concepts.
Reconstruction bottlenecks in object-centric generative models
Martin Engelcke, Oiwi Parker Jones, and Ingmar Posner · 2007
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2020
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
Causal world models by unsupervised deconfounding of physical dynamics, 2020
Minne Li, Mengyue Yang, Furui Liu, Xu Chen, Zhitang Chen, and Jun Wang · 2012
Earlier work this paper cites.
UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild, 2012
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation, 2014
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Convolutional LSTM network: A machine learning approach for precipitation nowcasting
Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Convolutional RNN: an Enhanced Model for Extracting Features from Sequential Data, 2017
Gil Keren and Björn Schuller · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Scrutinizing and de-biasing intuitive physics with neural stethoscopes, 2018
Fabian B. Fuchs, Oliver Groth, Adam R. Kosiorek, Alex Bewley, Markus Wulfmeier, Andrea Vedaldi, and Ingmar Posner · 2018
Earlier work this paper cites.
Shapestacks: Learning vision-based physical intuition for generalised object stacking
Oliver Groth, Fabian B Fuchs, Ingmar Posner, and Andrea Vedaldi · 2018
Earlier work this paper cites.
Compositional attention networks for machine reasoning
Drew A Hudson and Christopher D Manning · 2018
Earlier work this paper cites.
Embodied cognition
Peter König, Andrew Melnik, Caspar Goeke, Anna L Gert, Sabine U König, and Tim C Kietzmann · 2018
Earlier work this paper cites.
The world as an external memory: the price of saccades in a sensorimotor task
Andrew Melnik, Felix Schüler, Constantin A Rothkopf, and Peter König · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Earlier work this paper cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Earlier work this paper cites.
Relational deep reinforcement learning
Vinicius Zambaldi, David Raposo, Adam Santoro, Victor Bapst, Yujia Li, Igor Babuschkin, Karl Tuyls, David Reichert, Timothy Lillicrap, Edward Lockhart, et al · 2018
Earlier work this paper cites.
Distractor-aware siamese networks for visual object tracking
Zheng Zhu, Qiang Wang, Bo Li, Wei Wu, Junjie Yan, and Weiming Hu · 2018
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, et al · 2019
Earlier work this paper cites.
Phyre: A new benchmark for physical reasoning
Anton Bakhtin, Laurens van der Maaten, Justin Johnson, Laura Gustafson, and Ross Girshick · 2019
Earlier work this paper cites.
Cater: A diagnostic dataset for compositional actions and temporal reasoning
Rohit Girdhar and Deva Ramanan · 2019
Earlier work this paper cites.
Propagation networks for model-based control under partial observation
Yunzhu Li, Jiajun Wu, Jun-Yan Zhu, Joshua B Tenenbaum, Antonio Torralba, and Russ Tedrake · 2019
Earlier work this paper cites.
Ai meets angry birds
Jochen Renz, XiaoYu Ge, Matthew Stephenson, and Peng Zhang · 2019
Earlier work this paper cites.
Modeling expectation violation in intuitive physics with coarse probabilistic object representations
Kevin Smith, Lingjie Mei, Shunyu Yao, Jiajun Wu, Elizabeth Spelke, Josh Tenenbaum, and Tomer Ullman · 2019
Earlier work this paper cites.
Compositional video prediction
Yufei Ye, Maneesh Singh, Abhinav Gupta, and Shubham Tulsiani · 2019
Earlier work this paper cites.
Rapid trial-and-error learning with simulation supports flexible tool use and physical reasoning
Kelsey R Allen, Kevin A Smith, and Joshua B Tenenbaum · 2020
Earlier work this paper cites.
Craft: A benchmark for causal reasoning about forces and interactions
Tayfun Ates, M Samil Atesoglu, Cagatay Yigit, Ilker Kesen, Mert Kobas, Erkut Erdem, Aykut Erdem, Tilbe Goksun, and Deniz Yuret · 2020
Earlier work this paper cites.
An error-based addressing architecture for dynamic model learning
Nicolas Bach, Andrew Melnik, Federico Rosetto, and Helge Ritter · 2020
Earlier work this paper cites.
Learn to move through a combination of policy gradient algorithms: Ddpg, d4pg, and td3
Nicolas Bach, Andrew Melnik, Malte Schilling, Timo Korthals, and Helge Ritter · 2020
Earlier work this paper cites.
Compositional video synthesis with action graphs
Amir Bar, Roei Herzig, Xiaolong Wang, Anna Rohrbach, Gal Chechik, Trevor Darrell, and Amir Globerson · 2020
Earlier work this paper cites.
Cophy: Counterfactual learning of physical dynamics
Fabien Baradel, Natalia Neverova, Julien Mille, Greg Mori, and Christian Wolf · 2020
Cited alongside, same era.
Unsupervised discovery of 3d physical objects from video
Yilun Du, Kevin Smith, Tomer Ulman, Joshua Tenenbaum, and Jiajun Wu · 2020
Cited alongside, same era.
Relate: Physically plausible multi-object scene synthesis using structured latent spaces
Sebastien Ehrhardt, Oliver Groth, Aron Monszpart, Martin Engelcke, Ingmar Posner, Niloy Mitra, and Andrea Vedaldi · 2020
Cited alongside, same era.
Solving physics puzzles by reasoning about paths
Augustin Harter, Andrew Melnik, Gaurav Kumar, Dhruv Agarwal, Animesh Garg, and Helge Ritter · 2020
Cited alongside, same era.
Curl: Contrastive unsupervised representations for reinforcement learning
Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2020
Cited alongside, same era.
A survey of embodied ai: From simulators to research tasks
Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan · 2022
Later among the works it cites.
Towards improving the generation quality of autoregressive slot vaes, 2022
Patrick Emami, Pan He, Sanjay Ranka, and Anand Rangarajan · 2022
Later among the works it cites.
Solving reasoning tasks with a slot transformer
Ryan Faulkner and Daniel Zoran · 2022
Later among the works it cites.
Learning physical dynamics with subequivariant graph neural networks
Jiaqi Han, Wenbing Huang, Hengbo Ma, Jiachen Li, Josh Tenenbaum, and Chuang Gan · 2022
Later among the works it cites.
Make it move: controllable image-to-video generation with text descriptions
Yaosi Hu, Chong Luo, and Zhenzhong Chen · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical conditional relation networks for video question answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran · 2020
Cited alongside, same era.
Object-centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Cited alongside, same era.
Learning intuitive physics by explaining surprise
Hung Nguyen, Jay Patravali, Fuxin Li, and Alan Fern · 2020
Cited alongside, same era.
Learning Long-term Visual Dynamics with Region Proposal Interaction Networks
Haozhi Qi, Xiaolong Wang, Deepak Pathak, Yi Ma, and Jitendra Malik · 2020
Cited alongside, same era.
Ride: Rewarding impact-driven exploration for procedurally-generated environments
Roberta Raileanu and Tim Rocktäschel · 2020
Cited alongside, same era.
ESPRIT: Explaining solutions to physical reasoning tasks
Nazneen Fatema Rajani, Rui Zhang, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, and Dragomir Radev · 2020
Cited alongside, same era.
Learning object permanence from video
Aviv Shamsian, Ofri Kleinfeld, Amir Globerson, and Gal Chechik · 2020
Cited alongside, same era.
Steeven Janny, Fabien Baradel, Natalia Neverova, Madiha Nadri, Greg Mori, and Christian Wolf · 2022
Later among the works it cites.
Improving object-centric learning with query optimization, 2022
Baoxiong Jia, Yu Liu, and Siyuan Huang · 2022
Later among the works it cites.
Capturing temporal information in a single frame: Channel sampling strategies for action recognition
Kiyoon Kim, Shreyank N Gowda, Oisin Mac Aodha, and Laura Sevilla-Lara · 2022
Later among the works it cites.
Multimodal dialogue state tracking
Hung Le, Nancy F Chen, and Steven CH Hoi · 2022
Later among the works it cites.
Towards a unified neural architecture for visual recognition and reasoning
Calvin Luo, Ting Chen, Boqing Gong, and Chen Sun · 2022
Later among the works it cites.
Interactive language: Talking to robots in real time
Corey Lynch, Ayzaan Wahid, Jonathan Tompson, Tianli Ding, James Betker, Robert Baruch, Travis Armstrong, and Pete Florence · 2022
Later among the works it cites.
Behavioral cloning via search in video pretraining latent space
Federico Malato, Florian Leopold, Amogh Raut, Ville Hautamäki, and Andrew Melnik · 2022
Later among the works it cites.
A review of emerging research directions in abstract visual reasoning
Mikołaj Małkiński and Jacek Mańdziuk · 2022
Later among the works it cites.
Discovering generalizable spatial goal representations via graph-based active reward learning
Aviv Netanyahu, Tianmin Shu, Joshua Tenenbaum, and Pulkit Agrawal · 2022
Later among the works it cites.
CRIPP-VQA: Counterfactual reasoning about implicit physical properties via video question answering
Maitreya Patel, Tejas Gokhale, Chitta Baral, and Yezhou Yang · 2022
Later among the works it cites.
Intuitive physics learning in a deep-learning model inspired by developmental psychology
Luis S Piloto, Ari Weinstein, Peter Battaglia, and Matthew Botvinick · 2022
Later among the works it cites.
Learning temporal rules from noisy timeseries data
Karan Samel, Zelin Zhao, Binghong Chen, Shuang Li, Dharmashankar Subramanian, Irfan Essa, and Le Song · 2022
Later among the works it cites.
Towards learning implicit symbolic representation for visual reasoning
Chen Sun, Calvin Luo, Xingyi Zhou, Anurag Arnab, and Cordelia Schmid · 2022
Later among the works it cites.
Learning what and where: Disentangling location and identity tracking without supervision
Manuel Traub, Sebastian Otte, Tobias Menge, Matthias Karlbauer, Jannik Thuemmel, and Martin V Butz · 2022
Later among the works it cites.
Revealing occlusions with 4d neural fields
Basile Van Hoorick, Purva Tendulkar, Didac Suris, Dennis Park, Simon Stent, and Carl Vondrick · 2022
Later among the works it cites.
Slotformer: Unsupervised visual dynamics simulation with object-centric models
Ziyi Wu, Nikita Dvornik, Klaus Greff, Thomas Kipf, and Animesh Garg · 2022
Later among the works it cites.
Robotic skill acquisition via instruction augmentation with vision-language models
Ted Xiao, Harris Chan, Pierre Sermanet, Ayzaan Wahid, Anthony Brohan, Karol Hausman, Sergey Levine, and Jonathan Tompson · 2022
Later among the works it cites.
Tfcnet: Temporal fully connected networks for static unbiased temporal reasoning
Shiwen Zhang · 2022
Later among the works it cites.
A framework for the general design and computation of hybrid neural networks
Rong Zhao, Zheyu Yang, Hao Zheng, Yujie Wu, Faqiang Liu, Zhenzhi Wu, Lukai Li, Feng Chen, Seng Song, Jun Zhu, et al · 2022
Later among the works it cites.
Video question answering: Datasets, algorithms and challenges
Yaoyao Zhong, Junbin Xiao, Wei Ji, Yicong Li, Weihong Deng, and Tat-Seng Chua · 2022
Later among the works it cites.
Ddlp: Unsupervised object-centric video prediction with deep dynamic latent particles, 2023
Tal Daniel and Aviv Tamar · 2023
Closest in time.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al · 2023
Closest in time.
Probing neural representations of scene perception in a hippocampally dependent task using artificial neural networks
Markus Frey, Christian F Doeller, and Caswell Barry · 2023
Closest in time.
Vdt: An empirical study on video diffusion with transformers
Haoyu Lu, Guoxing Yang, Nanyi Fei, Yuqi Huo, Zhiwu Lu, Ping Luo, and Mingyu Ding · 2023
Closest in time.
Behavioral cloning via search in embedded demonstration dataset
Federico Malato, Florian Leopold, Ville Hautamaki, and Andrew Melnik · 2023
Closest in time.
Stephanie Milani, Anssi Kanervisto, Karolis Ramanauskas, Sander Schulhoff, Brandon Houghton, Sharada Mohanty, Byron Galbraith, Ke Chen, Yan Song, Tianze Zhou, et al · 2023
Closest in time.
Aran Nayebi, Rishi Rajalingham, Mehrdad Jazayeri, and Guangyu Robert Yang · 2023
Closest in time.
Contrastive language, action, and state pre-training for robot learning
Krishan Rana, Andrew Melnik, and Niko Sünderhauf · 2023
Closest in time.
Shape complexity estimation using vae
Markus Rothgaenger, Andrew Melnik, and Helge Ritter · 2023
Closest in time.
Language conditioned semantic search based policy for robotic manipulation tasks
Jannik Sheikh, Andrew Melnik, G Nandi, and Robert Haschke · 2023
Closest in time.
Intrinsic physical concepts discovery with object-centric predictive models
Qu Tang, Xiangyu Zhu, Zhen Lei, and Zhaoxiang Zhang · 2023
Closest in time.
Hsiao-Yu Tung, Mingyu Ding, Zhenfang Chen, Daniel Bear, Chuang Gan, Joshua B Tenenbaum, Daniel LK Yamins, Judith E Fan, and Kevin A Smith · 2023
Closest in time.
Perception and simulation during concept learning
Erik Weitnauer, Robert L Goldstone, and Helge Ritter · 2023
Closest in time.
Controllable video generation by learning the underlying dynamical system with neural ode
Yucheng Xu, Nanbo Li, Arushi Goel, Zijian Guo, Zonghai Yao, Hamidreza Kasaei, Mohammadreze Kasaei, and Zhibin Li · 2023
Closest in time.
Phy-q as a measure for physical reasoning intelligence
Cheng Xue, Vimukthini Pinto, Chathura Gamage, Ekaterina Nikonova, Peng Zhang, and Jochen Renz · 2023
Closest in time.
Probabilistic adaptation of text-to-video models
Mengjiao Yang, Yilun Du, Bo Dai, Dale Schuurmans, Joshua B Tenenbaum, and Pieter Abbeel · 2023
Closest in time.