Fetching the paper…
Reading the bibliography…
Strips: A new approach to the application of theorem proving to problem solving
R. E. Fikes and N. J. Nilsson · 1971
Earlier work this paper cites.
Pddl2. 1: An extension to pddl for expressing temporal planning domains
M. Fox and D. Long · 2003
Earlier work this paper cites.
Learning spatial relationships between objects
B. Rosman and S. Ramamoorthy · 2011
Earlier work this paper cites.
An industrial robotic knowledge representation for kit building applications
S. Balakirsky, Z. Kootbally, C. Schlenoff, T. Kramer, and S. Gupta · 2012
Earlier work this paper cites.
Learning spatial relationships from 3d vision using histograms
S. Fichtl, A. McManus, W. Mustafa, D. Kraft, N. Krüger, and F. Guerin · 2014
Earlier work this paper cites.
The ycb object and model set: Towards common benchmarks for manipulation research
B. Calli, A. Singh, A. Walsman, S. Srinivasa, P. Abbeel, and A. M. Dollar · 2015
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, et al · 2015
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Extended behavior trees for quick definition of flexible robotic tasks
F. Rovida, B. Grossmann, and V. Krüger · 2017
Earlier work this paper cites.
Goal-directed robot manipulation through axiomatic scene estimation
Z. Sui, L. Xiang, O. C. Jenkins, and K. Desingh · 2017
Earlier work this paper cites.
Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes
Y. Xiang, T. Schmidt, V. Narayanan, and D. Fox · 2017
Earlier work this paper cites.
Deep predictive policy training using reinforcement learning
A. Ghadirzadeh, A. Maki, D. Kragic, and M. Björkman · 2017
Earlier work this paper cites.
One-shot imitation learning
Y. Duan, M. Andrychowicz, B. C. Stadie, J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Implicit 3d orientation learning for 6d object detection from rgb images
M. Sundermeyer, Z.-C. Marton, M. Durner, M. Brucker, and R. Triebel · 2018
Cited alongside, same era.
Deepim: Deep iterative matching for 6d pose estimation
Y. Li, G. Wang, X. Ji, Y. Xiang, and D. Fox · 2018
Cited alongside, same era.
One-shot hierarchical imitation learning of compound visuomotor tasks
T. Yu, P. Abbeel, S. Levine, and C. Finn · 2018
Cited alongside, same era.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
K. Yi, J. Wu, C. Gan, A. Torralba, P. Kohli, and J. B. Tenenbaum · 2018
D. Ding, F. Hill, A. Santoro, and M. Botvinick · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Later among the works it cites.
ViSII: Virtual scene imaging interface, 2020
N. Morrical, J. Tremblay, S. Birchfield, and I. Wald · 2020
Later among the works it cites.
The epic-kitchens dataset: Collection, challenges and baselines
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2020
Later among the works it cites.
Poserbpf: A rao–blackwellized particle filter for 6-d object pose tracking
X. Deng, A. Mousavian, Y. Xiang, F. Xia, T. Bretl, and D. Fox · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Representing robot task plans as robust logical-dynamical systems
C. Paxton, N. Ratliff, C. Eppner, and D. Fox · 2019
Cited alongside, same era.
Neural task graphs: Generalizing to unseen tasks from a single video demonstration
D.-A. Huang, S. Nair, D. Xu, Y. Zhu, A. Garg, L. Fei-Fei, S. Savarese, and J. C. Niebles · 2019
Cited alongside, same era.
Prospection: Interpretable plans from language by predicting the future
C. Paxton, Y. Bisk, J. Thomason, A. Byravan, and D. Foxl · 2019
Cited alongside, same era.
The best of both modes: Separately leveraging rgb and depth for unseen object instance segmentation
C. Xie, Y. Xiang, A. Mousavian, and D. Fox · 2020
Cited alongside, same era.
Learning rgb-d feature embeddings for unseen object instance segmentation
Y. Xiang, C. Xie, A. Mousavian, and D. Fox · 2020
Cited alongside, same era.
Relational learning for skill preconditions
M. Sharma and O. Kroemer · 2020
Cited alongside, same era.
Transferable task execution from pixels through deep planning domain learning
K. Kase, C. Paxton, H. Mazhar, T. Ogata, and D. Fox · 2020
Cited alongside, same era.
Closest in time.
Grounding predicates through actions
T. Migimatsu and J. Bohg · 2021
Closest in time.
Hopper: Multi-hop transformer for spatiotemporal reasoning
H. Zhou, A. Kadav, F. Lai, A. Niculescu-Mizil, M. R. Min, M. Kapadia, and H. P. Graf · 2021
Closest in time.
Acronym: A large-scale grasp dataset based on simulation
C. Eppner, A. Mousavian, and D. Fox · 2021
Closest in time.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Closest in time.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2021
Closest in time.
Mdetr–modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, I. Misra, G. Synnaeve, and N. Carion · 2021
Closest in time.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Closest in time.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Closest in time.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Closest in time.