Fetching the paper…
Reading the bibliography…
Our goal is to generate a policy to complete an unseen task given just a single video demonstration of the task in a given domain.
Strips: A new approach to the application of theorem proving to problem solving
R. E. Fikes and N. J. Nilsson · 1971
Earlier work this paper cites.
Learning and executing generalized robot plans
R. E. Fikes, P. E. Hart, and N. J. Nilsson · 1972
Earlier work this paper cites.
A structure for plans and behavior
E. D. Sacerdoti · 1975
Earlier work this paper cites.
A robust layered control system for a mobile robot
R. Brooks · 1986
Earlier work this paper cites.
Shop: Simple hierarchical ordered planner
D. Nau, Y. Cao, A. Lotem, and H. Munoz-Avila · 1999
Earlier work this paper cites.
A hierarchical architecture for behavior-based robots
M. N. Nicolescu and M. J. Matarić · 2002
Earlier work this paper cites.
Understanding videos, constructing plots learning a visually grounded storyline model from annotated videos
A. Gupta, P. Srinivasan, J. Shi, and L. S. Davis · 2009
Earlier work this paper cites.
Towards one shot learning by imitation for humanoid robots
Y. Wu and Y. Demiris · 2010
Earlier work this paper cites.
Hierarchical task and motion planning in the now
L. P. Kaelbling and T. Lozano-Pérez · 2011
Earlier work this paper cites.
Keyframe-based learning from demonstration
B. Akgun, M. Cakmak, K. Jiang, and A. L. Thomaz · 2012
Earlier work this paper cites.
Learning and generalization of complex tasks from unstructured demonstrations
S. Niekum, S. Osentoski, G. Konidaris, and A. G. Barto · 2012
Earlier work this paper cites.
Jhu-isi gesture and skill assessment working set ( jigsaws ) : A surgical activity dataset for human motion modeling
Y. Gao, S. S. Vedula, C. E. Reiley, N. Ahmidi, B. Varadarajan, H. C. Lin, L. Tao, L. Zappella, B. Béjar, D. D. Yuh, C. C. G. Chen, R. Vidal, S. Khudanpur, and G. D. Hager · 2014
Earlier work this paper cites.
Combined task and motion planning through an extensible planner-independent interface layer
S. Srivastava, E. Fang, L. Riano, R. Chitnis, S. Russell, and P. Abbeel · 2014
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
M.-T. Luong, H. Pham, and C. D. Manning · 2015
Earlier work this paper cites.
Unsupervised semantic parsing of video collections
O. Sener, A. R. Zamir, S. Savarese, and A. Saxena · 2015
Earlier work this paper cites.
Book2movie: Aligning video scenes with book chapters
M. Tapaswi, M. Bauml, and R. Stiefelhagen · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Cited alongside, same era.
Unsupervised learning from narrated instruction videos
J.-B. Alayrac, P. Bojanowski, N. Agrawal, I. Laptev, J. Sivic, and S. Lacoste-Julien · 2016
Cited alongside, same era.
Deep compositional question answering with neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Autonomously constructing hierarchical task networks for planning and human-robot collaboration
B. Hayes and B. Scassellati · 2016
Cited alongside, same era.
Jointly learning grounded task structures from language instruction and visual demonstration
C. Liu, S. Yang, S. Saba-Sadiya, N. Shukla, Y. He, S.-C. Zhu, and J. Chai · 2016
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
T. N. Kipf and M. Welling · 2017
Later among the works it cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
E. Kolve, R. Mottaghi, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi · 2017
Later among the works it cites.
Dense-captioning events in videos
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles · 2017
Later among the works it cites.
Time-contrastive networks: Self-supervised learning from multi-view observation
P. Sermanet, C. Lynch, J. Hsu, and S. Levine · 2017
Later among the works it cites.
Comparing Robot Grasping Teleoperation across Desktop and Virtual Reality with ROS Reality
D. Whitney, E. Rosen, E. Phillips, G. Konidaris, and S. Tellex · 2017
Later among the works it cites.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The curious robot: Learning visual representations via physical interactions
L. Pinto, D. Gandhi, Y. Han, Y.-L. Park, and A. Gupta · 2016
Cited alongside, same era.
Modular multitask reinforcement learning with policy sketches
J. Andreas, D. Klein, and S. Levine · 2017
Cited alongside, same era.
pybullet, a python module for physics simulation, games, robotics and machine learning
E. Coumans and Y. Bai · 2017
Cited alongside, same era.
Learning Modular Neural Network Policies for Multi-Task and Multi-Robot Transfer
C. Devin, A. Gupta, T. Darrell, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
One-Shot Imitation Learning
Y. Duan, M. Andrychowicz, B. C. Stadie, J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
One-shot visual imitation learning via meta-learning
C. Finn, T. Yu, T. Zhang, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Cognitive mapping and planning for visual navigation
S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik · 2017
Cited alongside, same era.
T. Zhang, Z. McCarthy, O. Jow, D. Lee, K. Goldberg, and P. Abbeel · 2017
Later among the works it cites.
Visual semantic planning using deep successor representations
Y. Zhu, D. Gordon, E. Kolve, D. Fox, L. Fei-Fei, A. Gupta, R. Mottaghi, and A. Farhadi · 2017
Later among the works it cites.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Y. Zhu, R. Mottaghi, E. Kolve, J. J. Lim, A. Gupta, L. Fei-Fei, and A. Farhadi · 2017
Later among the works it cites.
One-shot learning of multi-step tasks from observation via activity localization in auxiliary video
W. Goo and S. Niekum · 2018
Closest in time.
Finding “it”: Weakly-supervised reference-aware visual grounding in instructional videos
D.-A. Huang, S. Buch, L. Dery, A. Garg, L. Fei-Fei, and J. C. Niebles · 2018
Closest in time.
Visual coreference resolution in visual dialog using neural module networks
S. Kottur, J. M. Moura, D. Parikh, D. Batra, and M. Rohrbach · 2018
Closest in time.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
Y. Liu, A. Gupta, P. Abbeel, and S. Levine · 2018
Closest in time.
Nervenet: Learning structured policy with graph neural networks
T. Wang, R. Liao, J. Ba, and S. Fidler · 2018
Closest in time.
Neural task programming: Learning to generalize across hierarchical tasks
D. Xu, S. Nair, Y. Zhu, J. Gao, A. Garg, L. Fei-Fei, and S. Savarese · 2018
Closest in time.
One-shot imitation from observing humans via domain-adaptive meta-learning
T. Yu, C. Finn, A. Xie, S. Dasari, T. Zhang, P. Abbeel, and S. Levine · 2018
Closest in time.