Fetching the paper…
Reading the bibliography…
In our pursuit of advancing multi-modal AI assistants capable of guiding users to achieve complex multi-step goals, we propose the task of "Visual Planning for Assistance (VPA)".
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
A hierarchical representation for future action prediction
Tian Lan, Tsung-Chuan Chen, and Silvio Savarese · 2014
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
William Lotter, Gabriel Kreiman, and David Cox · 2016
Earlier work this paper cites.
Anticipating visual representations from unlabeled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
First-person action-object detection with egonet
Gedas Bertasius, Hyun Soo Park, X Yu Stella, and Jianbo Shi · 2017
Earlier work this paper cites.
Next-active-object prediction from egocentric videos
Antonino Furnari, Sebastiano Battiato, Kristen Grauman, and Giovanni Maria Farinella · 2017
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Mel Vecerik, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin Riedmiller · 2017
Earlier work this paper cites.
Decomposing motion and content for natural video sequence prediction
Ruben Villegas, Jimei Yang, Seunghoon Hong, Xunyu Lin, and Honglak Lee · 2017
Earlier work this paper cites.
When will you do what?-anticipating temporal occurrences of activities
Yazan Abu Farha, Alexander Richard, and Juergen Gall · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2018
Earlier work this paper cites.
Deep q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, et al · 2018
Earlier work this paper cites.
Policy optimization with demonstrations
Bingyi Kang, Zequn Jie, and Jiashi Feng · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Earlier work this paper cites.
What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention
Antonino Furnari and Giovanni Maria Farinella · 2019
Earlier work this paper cites.
Two body problem: Collaborative visual task completion
Unnat Jain, Luca Weihs, Eric Kolve, Mohammad Rastegari, Svetlana Lazebnik, Ali Farhadi, Alexander G. Schwing, and Aniruddha Kembhavi · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
A variational auto-encoder model for stochastic point processes. in 2019 ieee
Nazanin Mehrasa, Akash Abdu Jyothi, Thibaut Durand, Jiawei He, Leonid Sigal, and Greg Mori · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Habitat: A Platform for Embodied AI Research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra · 2019
Earlier work this paper cites.
Zero-shot anticipation for instructional activities
Fadime Sener and Angela Yao · 2019
Earlier work this paper cites.
Coin: A large-scale dataset for comprehensive instructional video analysis
Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou · 2019
Earlier work this paper cites.
Cross-task weakly supervised learning from instructional videos
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Procedure planning in instructional videos
Chien-Yi Chang, De-An Huang, Danfei Xu, Ehsan Adeli, Li Fei-Fei, and Juan Carlos Niebles · 2020
Earlier work this paper cites.
Soundspaces: Audio-visual navigation in 3d environments
Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip Robinson, and Kristen Grauman · 2020
Earlier work this paper cites.
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2020
Earlier work this paper cites.
Long-term anticipation of activities with cycle consistency
Yazan Abu Farha, Qiuhong Ke, Bernt Schiele, and Juergen Gall · 2020
Cited alongside, same era.
A cordial sync: Going beyond marginal policies for multi-agent embodied tasks
Unnat Jain, Luca Weihs, Eric Kolve, Ali Farhadi, Svetlana Lazebnik, Aniruddha Kembhavi, and Alexander G. Schwing · 2020
Cited alongside, same era.
Ego-topo: Environment affordances from egocentric video
Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman · 2020
Cited alongside, same era.
Temporal aggregate representations for long-range video understanding
Fadime Sener, Dipika Singhania, and Angela Yao · 2020
Cited alongside, same era.
Understanding human hands in contact at internet scale
Dandan Shan, Jiaqi Geng, Michelle Shu, and David Fouhey · 2020
Cited alongside, same era.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch · 2022
Later among the works it cites.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al · 2022
Later among the works it cites.
Vima: General robot manipulation with multimodal prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan · 2022
Later among the works it cites.
Universal Approximation Under Constraints is Possible with Transformers, Feb. 2022
Anastasis Kratsios, Behnoosh Zamanlooy, Tianlin Liu, and Ivan Dokmanić · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saim Wani, Shivansh Patel, Unnat Jain, Angel Chang, and Manolis Savva · 2020
Cited alongside, same era.
Allenact: A framework for embodied ai research
Luca Weihs, Jordi Salvador, Klemen Kotar, Unnat Jain, Kuo-Hao Zeng, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2020
Cited alongside, same era.
Procedure planning in instructional videos via contextual modeling and model-based policy learning
Jing Bi, Jiebo Luo, and Chenliang Xu · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Cited alongside, same era.
Anticipative video transformer
Rohit Girdhar and Kristen Grauman · 2021
Cited alongside, same era.
Gridtopix: Training embodied agents with minimal supervision
Unnat Jain, Iou-Jen Liu, Svetlana Lazebnik, Aniruddha Kembhavi, Luca Weihs, and Alexander G Schwing · 2021
Cited alongside, same era.
Shuang Li, Xavier Puig, Yilun Du, Clinton Wang, Ekin Akyurek, Antonio Torralba, Jacob Andreas, and Igor Mordatch · 2022
Later among the works it cites.
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng · 2022
Later among the works it cites.
Joint hand motion and interaction hotspots prediction from egocentric videos
Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, and Xiaolong Wang · 2022
Later among the works it cites.
Long-term action forecasting using multi-headed attention-based variational recurrent neural networks
Siyuan Brandon Loh, Debaditya Roy, and Basura Fernando · 2022
Later among the works it cites.
Unified-io: A unified model for vision, language, and multi-modal tasks
Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2022
Later among the works it cites.
Learning state-aware visual representations from audible interactions
Himangi Mittal, Pedro Morgado, Unnat Jain, and Abhinav Gupta · 2022
Later among the works it cites.
History compression via language models in reinforcement learning
Fabian Paischer, Thomas Adler, Vihang Patil, Angela Bitto-Nemling, Markus Holzleitner, Sebastian Lehner, Hamid Eghbal-Zadeh, and Sepp Hochreiter · 2022
Later among the works it cites.
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al · 2022
Later among the works it cites.
Transferring knowledge from text to video: Zero-shot anticipation for procedural actions
Fadime Sener, Rishabh Saraf, and Angela Yao · 2022
Later among the works it cites.
Progprompt: Generating situated robot task plans using large language models
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg · 2022
Later among the works it cites.
Behavior: Benchmark for everyday household activities in virtual, interactive, and ecological environments
Sanjana Srivastava, Chengshu Li, Michael Lingelbach, Roberto Martín-Martín, Fei Xia, Kent Elliott Vainio, Zheng Lian, Cem Gokmen, Shyamal Buch, Karen Liu, et al · 2022
Later among the works it cites.
Plate: Visually-grounded planning with transformers in procedural tasks
Jiankai Sun, De-An Huang, Bo Lu, Yun-Hui Liu, Bolei Zhou, and Animesh Garg · 2022
Later among the works it cites.
Are Transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2022
Later among the works it cites.
P3iv: Probabilistic procedure planning from instructional videos with weak supervision
He Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G Derpanis, Richard P Wildes, and Allan D Jepson · 2022
Later among the works it cites.
Affordances from human videos as a versatile representation for robotics
Shikhar Bahl, Russell Mendonca, Lili Chen, Unnat Jain, and Deepak Pathak · 2023
Closest in time.
Collaborating with language models for embodied reasoning
Ishita Dasgupta, Christine Kaeser-Chen, Kenneth Marino, Arun Ahuja, Sheila Babayan, Felix Hill, and Rob Fergus · 2023
Closest in time.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence · 2023
Closest in time.
Language is not all you need: Aligning perception with language models
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, et al · 2023
Closest in time.
Action dynamics task graphs for learning plannable representations of procedural tasks
Weichao Mao, Ruta Desai, Michael Louis Iuzzolino, and Nitin Kamra · 2023
Closest in time.
Intention-conditioned long-term human egocentric action anticipation
Esteve Valls Mascaró, Hyemin Ahn, and Dongheui Lee · 2023
Closest in time.
Intention-conditioned long-term human egocentric action anticipation
Esteve Valls Mascaro, Hyemin Ahn, and Dongheui Lee · 2023
Closest in time.
Structured world models from human vidoes
Russell Mendonca, Shikhar Bahl, and Deepak Pathak · 2023
Closest in time.
Adaptive coordination for social embodied rearrangement
Andrew Szot, Unnat Jain, Dhruv Batra, Zsolt Kira, Ruta Desai, and Akshara Rai · 2023
Closest in time.
Ego-only: Egocentric action detection without exocentric pretraining
Huiyu Wang, Mitesh Kumar Singh, and Lorenzo Torresani · 2023
Closest in time.
Exif as language: Learning cross-modal associations between images and camera metadata
Chenhao Zheng, Ayush Shrivastava, and Andrew Owens · 2023
Closest in time.