Fetching the paper…
Reading the bibliography…
Transformers have revolutionized vision and natural language processing with their ability to scale with large datasets.
Building an environment model using depth information
Y. Roth-Tabak and R. Jain · 1989
Earlier work this paper cites.
New approaches to robotics
R. A. Brooks · 1991
Earlier work this paper cites.
Robot spatial perceptionby stereoscopic vision and 3d evidence grids
H. Moravec · 1996
Earlier work this paper cites.
A tutorial on energy-based learning
Y. LeCun, S. Chopra, R. Hadsell, M. Ranzato, and F. Huang · 2006
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
S. Tellex, T. Kollar, S. Dickerson, M. Walter, A. Banerjee, S. Teller, and N. Roy · 2011
Earlier work this paper cites.
Interpreting and executing recipes with a cooking robot
M. Bollini, S. Tellex, T. Thompson, N. Roy, and D. Rus · 2013
Earlier work this paper cites.
V-rep: A versatile and scalable robot simulation framework
E. Rohmer, S. P. N. Singh, and M. Freese · 2013
Earlier work this paper cites.
Integrated task and motion planning in belief space
L. P. Kaelbling and T. Lozano-Pérez · 2013
Earlier work this paper cites.
Real-time 3d reconstruction at scale using voxel hashing
M. Nießner, M. Zollhöfer, S. Izadi, and M. Stamminger · 2013
Earlier work this paper cites.
The ecological approach to visual perception: classic edition
J. J. Gibson · 2014
Earlier work this paper cites.
Single image 3d object detection and pose estimation for grasping
M. Zhu, K. G. Derpanis, Y. Yang, S. Brahmbhatt, M. Zhang, C. Phillips, M. Lecce, and K. Daniilidis · 2014
Earlier work this paper cites.
Learning from unscripted deictic gesture and language for human-robot interactions
C. Matuszek, L. Bo, L. Zettlemoyer, and D. Fox · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Learning to interpret natural language commands through human-robot dialog
J. Thomason, S. Zhang, R. J. Mooney, and P. Stone · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Tell me dave: Context-sensitive grounding of natural language to manipulation instructions
D. K. Misra, J. Sung, K. Lee, and A. Saxena · 2016
Earlier work this paper cites.
Natural language communication with robots
Y. Bisk, D. Yuret, and D. Marcu · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Multi-view self-supervised deep learning for 6d pose estimation in the amazon picking challenge
A. Zeng, K.-T. Yu, S. Song, D. Suo, E. Walker, A. Rodriguez, and J. Xiao · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes
Y. Xiang, T. Schmidt, V. Narayanan, and D. Fox · 2018
Earlier work this paper cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Interactive visual grounding of referring expressions for human-robot interaction
M. Shridhar and D. Hsu · 2018
Earlier work this paper cites.
Interactively picking real-world objects with unconstrained spoken language instructions
J. Hatori, Y. Kikuchi, S. Kobayashi, K. Takahashi, Y. Tsuboi, Y. Unno, W. Ko, and J. Tan · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Earlier work this paper cites.
Mapping instructions to actions in 3d environments with visual goal prediction
D. Misra, A. Bennett, V. Blukis, E. Niklasson, M. Shatkhin, and Y. Artzi · 2018
Earlier work this paper cites.
From skills to symbols: Learning symbolic representations for abstract high-level planning
G. Konidaris, L. P. Kaelbling, and T. Lozano-Perez · 2018
Earlier work this paper cites.
Alphastar: Mastering the real-time strategy game starcraft ii
O. Vinyals, I. Babuschkin, J. Chung, M. Mathieu, M. Jaderberg, W. M. Czarnecki, A. Dudzik, A. Huang, P. Georgiev, R. Powell, et al · 2019
Earlier work this paper cites.
Robotic pick-and-place of novel objects in clutter with multi-affordance grasping and cross-domain image matching
A. Zeng, S. Song, K.-T. Yu, E. Donlon, F. R. Hogan, M. Bauza, D. Ma, O. Taylor, M. Liu, E. Romo, et al · 2019
Earlier work this paper cites.
6-dof graspnet: Variational grasp generation for object manipulation
A. Mousavian, C. Eppner, and D. Fox · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
Prospection: Interpretable plans from language by predicting the future
C. Paxton, Y. Bisk, J. Thomason, A. Byravan, and D. Fox · 2019
Earlier work this paper cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh · 2019
Earlier work this paper cites.
Pyrep: Bringing v-rep to deep robot learning
S. James, M. Freese, and A. J. Davison · 2019
Earlier work this paper cites.
Deepvoxels: Learning persistent 3d feature embeddings
V. Sitzmann, J. Thies, F. Heide, M. Nießner, G. Wetzstein, and M. Zollhofer · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Cited alongside, same era.
Transporter networks: Rearranging the visual world for robotic manipulation
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, and J. Lee · 2020
Cited alongside, same era.
Self-supervised 6d object pose estimation for robot manipulation
X. Deng, Y. Xiang, A. Mousavian, C. Eppner, T. Bretl, and D. Fox · 2020
Cited alongside, same era.
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard · 2021
Later among the works it cites.
Coarse-to-fine imitation learning: Robot manipulation from a single demonstration
E. Johns · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Later among the works it cites.
Multi-task reinforcement learning with context-based representations
S. Sodhani, A. Zhang, and J. Pineau · 2021
Later among the works it cites.
ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
M. Shridhar, X. Yuan, M.-A. Côté, Y. Bisk, A. Trischler, and M. Hausknecht · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The best of both modes: Separately leveraging rgb and depth for unseen object instance segmentation
C. Xie, Y. Xiang, A. Mousavian, and D. Fox · 2020
Cited alongside, same era.
Learning to Manipulate Deformable Objects without Demonstrations
Y. Wu, W. Yan, T. Kurutach, L. Pinto, and P. Abbeel · 2020
Cited alongside, same era.
Grasping in the wild: Learning 6dof closed-loop grasping from low-cost demonstrations
S. Song, A. Zeng, J. Lee, and T. Funkhouser · 2020
Cited alongside, same era.
6-dof grasping for target-driven object manipulation in clutter
A. Murali, A. Mousavian, C. Eppner, C. Paxton, and D. Fox · 2020
Cited alongside, same era.
Transformers for one-shot visual imitation
S. Dasari and A. Gupta · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Cited alongside, same era.
Few-shot object grounding for mapping natural language instructions to robot control
V. Blukis, R. A. Knepper, and Y. Artzi · 2020
Cited alongside, same era.
Integrated task and motion planning
C. R. Garrett, R. Chitnis, R. Holladay, B. Kim, T. Silver, L. P. Kaelbling, and T. Lozano-Pérez · 2021
Later among the works it cites.
Voxel transformer for 3d object detection
J. Mao, Y. Xue, M. Niu, H. Bai, J. Feng, X. Liang, H. Xu, and C. Xu · 2021
Later among the works it cites.
Coconets: Continuous contrastive 3d scene representations
S. Lal, M. Prabhudesai, I. Mediratta, A. W. Harley, and K. Fragkiadaki · 2021
Later among the works it cites.
Coordination among neural modules through a shared global workspace
A. Goyal, A. R. Didolkar, A. Lamb, K. Badola, N. R. Ke, N. Rahaman, J. Binas, C. Blundell, M. C. Mozer, and Y. Bengio · 2021
Later among the works it cites.
Vision transformers for dense prediction
R. Ranftl, A. Bochkovskiy, and V. Koltun · 2021
Later among the works it cites.
Mdetr–modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, I. Misra, G. Synnaeve, and N. Carion · 2021
Later among the works it cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
A. Birhane, V. U. Prabhu, and E. Kahembwe · 2021
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Later among the works it cites.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al · 2022
Closest in time.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, et al · 2022
Closest in time.
Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation
S. James, K. Wada, T. Laidlow, and A. J. Davison · 2022
Closest in time.
Guiding multi-step rearrangement tasks with natural language instructions
E. Stengel-Eskin, A. Hundt, Z. He, A. Murali, N. Gopalan, M. Gombolay, and G. Hager · 2022
Closest in time.
Umpnet: Universal manipulation policy network for articulated objects
Z. Xu, H. Zhanpeng, and S. Song · 2022
Closest in time.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Closest in time.
Multi-game decision transformers
K.-H. Lee, O. Nachum, M. Yang, L. Lee, D. Freeman, W. Xu, S. Guadarrama, I. Fischer, E. Jang, H. Michalewski, et al · 2022
Closest in time.
Metamorph: Learning universal controllers with transformers
A. Gupta, L. Fan, S. Ganguli, and L. Fei-Fei · 2022
Closest in time.
Structformer: Learning spatial structure for language-guided semantic rearrangement of novel objects
W. Liu, C. Paxton, T. Hermans, and D. Fox · 2022
Closest in time.
Learning language-conditioned robot behavior from offline data and crowd-sourced annotation
S. Nair, E. Mitchell, K. Chen, S. Savarese, C. Finn, et al · 2022
Closest in time.
What matters in language conditioned robotic imitation learning over unstructured data
O. Mees, L. Hermann, and W. Burgard · 2022
Closest in time.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
S. James and A. J. Davison · 2022
Closest in time.
Auto-lambda: Disentangling dynamic task relationships
S. Liu, S. James, A. J. Davison, and E. Johns · 2022
Closest in time.
On the effectiveness of fine-tuning versus meta-reinforcement learning
Z. Mandi, P. Abbeel, and S. James · 2022
Closest in time.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, A. Wong, S. Welker, K. Choromanski, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, et al · 2022
Closest in time.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Closest in time.
Voxel set transformer: A set-to-set approach to 3d object detection from point clouds
C. He, R. Li, S. Li, and L. Zhang · 2022
Closest in time.
Towards optimal correlational object search
K. Zheng, R. Chitnis, Y. Sung, G. Konidaris, and S. Tellex · 2022
Closest in time.
A persistent spatial semantic representation for high-level natural language instruction execution
V. Blukis, C. Paxton, D. Fox, A. Garg, and Y. Artzi · 2022
Closest in time.
Voxel-informed language grounding
R. Corona, S. Zhu, D. Klein, and T. Darrell · 2022
Closest in time.
Instant neural graphics primitives with a multiresolution hash encoding
T. Müller, A. Evans, C. Schied, and A. Keller · 2022
Closest in time.
Plenoxels: Radiance fields without neural networks
Sara Fridovich-Keil and Alex Yu, M. Tancik, Q. Chen, B. Recht, and A. Kanazawa · 2022
Closest in time.
Coarse-to-fine q-attention with learned path ranking
S. James and P. Abbeel · 2022
Closest in time.
Implicit behavioral cloning
P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson · 2022
Closest in time.