Fetching the paper…
Reading the bibliography…
In human environments, robots are expected to accomplish a variety of manipulation tasks given simple natural language instructions.
Real time control of a robot with a mobile camera
J. Hill · 1979
Earlier work this paper cites.
Learning many related tasks at the same time with backpropagation
R. Caruana · 1994
Earlier work this paper cites.
Discovering structure in multiple learning tasks: The TC algorithm
S. Thrun and J. O’Sullivan · 1996
Earlier work this paper cites.
Reinforcement learning for humanoid robotics
J. Peters, S. Vijayakumar, and S. Schaal · 2003
Earlier work this paper cites.
Visual servo control. i. basic approaches
F. Chaumette and S. Hutchinson · 2006
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
S. Tellex, T. Kollar, S. Dickerson, M. R. Walter, A. G. Banerjee, S. Teller, and N. Roy · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Apriltag: A robust and flexible visual fiducial system
E. Olson · 2011
Earlier work this paper cites.
Online multi-task learning for policy gradient methods
H. B. Ammar, E. Eaton, P. Ruvolo, and M. Taylor · 2014
Earlier work this paper cites.
UNet: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Efficient grounding of abstract spatial concepts for natural language interaction with robot manipulators
R. Paul, J. Arkin, N. Roy, and T. M Howard · 2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Learning visual servoing with deep features and fitted q-iteration
A. X. Lee, S. Levine, and P. Abbeel · 2017
Earlier work this paper cites.
Learning modular neural network policies for multi-task and multi-robot transfer
C. Devin, A. Gupta, T. Darrell, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Earlier work this paper cites.
One-shot visual imitation learning via meta-learning
C. Finn, T. Yu, T. Zhang, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Mapping instructions and visual observations to actions with reinforcement learning
D. Misra, J. Langford, and Y. Artzi · 2017
Earlier work this paper cites.
Grounded language learning in a simulated 3D world
K. M. Hermann, F. Hill, S. Green, F. Wang, R. Faulkner, H. Soyer, D. Szepesvari, W. M. Czarnecki, M. Jaderberg, D. Teplyashin, M. Wainwright, C. Apps, D. Hassabis, and P. Blunso · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
FiLM: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Earlier work this paper cites.
Training deep neural networks for visual servoing
Q. Bateux, E. Marchand, J. Leitner, F. Chaumette, and P. Corke · 2018
Earlier work this paper cites.
One-shot imitation from observing humans via domain-adaptive meta-learning
T. Yu, C. Finn, A. Xie, S. Dasari, T. Zhang, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, and P. Abbeel · 2018
Earlier work this paper cites.
Modulated policy hierarchies
A. Pashevich, D. Hafner, J. Davidson, R. Sukthankar, and C. Schmid · 2018
Earlier work this paper cites.
Behavioral cloning from observation
F. Torabi, G. Warnell, and P. Stone · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Cited alongside, same era.
Solving rubik’s cube with a robot hand
I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, et al · 2019
Cited alongside, same era.
Disentangled relational representations for explaining and learning from demonstration
Y. Hristov, D. Angelov, M. Burke, A. Lascarides, and S. Ramamoorthy · 2019
Cited alongside, same era.
Multi-task reinforcement learning with context-based representations
S. Sodhani, A. Zhang, and J. Pineau · 2021
Later among the works it cites.
MT-Opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Later among the works it cites.
Learning generalizable robotic reward functions from ”in-the-wild” human videos
A. S. Chen, S. Nair, and C. Finn · 2021
Later among the works it cites.
ManipulaTHOR: A framework for visual object manipulation
K. Ehsani, W. Han, A. Herrasti, E. VanderBilt, L. Weihs, E. Kolve, A. Kembhavi, and R. Mottaghi · 2021
Later among the works it cites.
Guiding multi-step rearrangement tasks with natural language instructions
E. Stengel-Eskin, A. Hundt, Z. He, A. Murali, N. Gopalan, M. Gombolay, and G. Hager · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lxmert: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Cited alongside, same era.
Multilingual universal sentence encoder for semantic retrieval
Y. Yang, D. Cer, A. Ahmad, M. Guo, J. Law, N. Constant, G. H. Abrego, S. Yuan, C. Tar, Y.-H. Sung, et al · 2020
Cited alongside, same era.
RLBench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Cited alongside, same era.
Learning to combine primitive skills: A step towards versatile robotic manipulation
R. Strudel, A. Pashevich, I. Kalevatykh, I. Laptev, and C. Schmid · 2020
Cited alongside, same era.
Deep learning in robotics: Survey on model structures and training strategies
A. I. Károly, P. Galambos, J. Kuti, and I. J. Rudas · 2020
Cited alongside, same era.
Cosypose: Consistent multi-view multi-object 6d pose estimation
Y. Labbe, J. Carpentier, M. Aubry, and J. Sivic · 2020
Cited alongside, same era.
D. I. A. Team, J. Abramson, A. Ahuja, A. Brussee, F. Carnevale, M. Cassin, F. Fischer, P. Georgiev, A. Goldin, T. Harley, et al · 2021
Later among the works it cites.
Semantically grounded object matching for robust robotic scene rearrangement
W. Goodwin, S. Vaze, I. Havoutis, and I. Posner · 2021
Later among the works it cites.
StructFormer: Learning spatial structure for language-guided semantic rearrangement of novel objects
W. Liu, C. Paxton, T. Hermans, and D. Fox · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Later among the works it cites.
Pretraining for language conditioned imitation with transformers
A. L. Putterman, K. Lu, I. Mordatch, and P. Abbeel · 2021
Later among the works it cites.
History aware multimodal transformer for vision-and-language navigation
S. Chen, P.-L. Guhur, C. Schmid, and I. Laptev · 2021
Later among the works it cites.
Episodic transformer for vision-and-language navigation
A. Pashevich, C. Schmid, and C. Sun · 2021
Later among the works it cites.
Airbert: In-domain pretraining for vision-and-language navigation
P.-L. Guhur, M. Tapaswi, S. Chen, I. Laptev, and C. Schmid · 2021
Later among the works it cites.
Lancon-learn: Learning with language to enable generalization in multi-task manipulation
A. Silva, N. Moorman, W. Silva, Z. Zaidi, N. Gopalan, and M. Gombolay · 2021
Later among the works it cites.
CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard · 2022
Closest in time.
CLIPort: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Closest in time.
BC-Z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Closest in time.
Q-Attention: Enabling efficient learning for vision-based robotic manipulation
S. James and A. J. Davison · 2022
Closest in time.
Coarse-to-fine Q-Attention: Efficient learning for visual robotic manipulation via discretisation
S. James, K. Wada, T. Laidlow, and A. J. Davison · 2022
Closest in time.
Auto-Lambda: Disentangling dynamic task relationships
S. Liu, S. James, A. J. Davison, and E. Johns · 2022
Closest in time.
Learning language-conditioned robot behavior from offline data and crowd-sourced annotation
S. Nair, E. Mitchell, K. Chen, S. Savarese, C. Finn, et al · 2022
Closest in time.
Implicit behavioral cloning
P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson · 2022
Closest in time.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Closest in time.
LISA: Learning interpretable skill abstractions from language
D. Garg, S. Vaidyanath, K. Kim, J. Song, and S. Ermon · 2022
Closest in time.
Simple but effective: Clip embeddings for embodied ai
A. Khandelwal, L. Weihs, R. Mottaghi, and A. Kembhavi · 2022
Closest in time.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al · 2022
Closest in time.
Coarse-to-fine Q-Attention with learned path ranking
S. James and P. Abbeel · 2022
Closest in time.