Fetching the paper…
Reading the bibliography…
Despite tremendous advances in AI, it remains a significant challenge to develop interactive task guidance systems that can offer situated, personalized guidance and assist humans in various tasks.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
A Coefficient of Agreement for Nominal Scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
Preliminary investigation of wearable computers for task guidance in aircraft inspection
Jennifer J Ockerman and Amy R Pritchett. 1998 · 1998
Earlier work this paper cites.
A review and reappraisal of task guidance: Aiding workers in procedure following
Jennifer Ockerman and Amy Pritchett. 2000 · 2000
Earlier work this paper cites.
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. 2020 · 2006
Earlier work this paper cites.
Social interactions: A first-person perspective
Alircza Fathi, Jessica K Hodgins, and James M Rehg. 2012 · 2012
Earlier work this paper cites.
Discovering important people and objects for egocentric video summarization
Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman. 2012 · 2012
Earlier work this paper cites.
Detecting activities of daily living in first-person camera views
Hamed Pirsiavash and Deva Ramanan. 2012 · 2012
Earlier work this paper cites.
Physical causality of action verbs in grounded language understanding
Qiaozi Gao, Malcolm Doering, Shaohua Yang, and Joyce Chai. 2016 · 2016
Earlier work this paper cites.
Detecting engagement in egocentric video
Yu-Chuan Su and Kristen Grauman. 2016 · 2016
Earlier work this paper cites.
Multi-modal augmented-reality assembly guidance based on bare-hand interface
Xuan Wang, SK Ong, and Andrew Yeh-Ching Nee. 2016 · 2016
Earlier work this paper cites.
Automated capture and delivery of assistive task guidance with an eyewear computer: the glaciar system
Teesid Leelasawassuk, Dima Damen, and Walterio Mayol-Cuevas. 2017 · 2017
Earlier work this paper cites.
Procnets: Learning to segment procedures in untrimmed and unconstrained videos
Luowei Zhou, Chenliang Xu, and Jason J. Corso. 2017 · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. 2018 · 2018
Earlier work this paper cites.
In the eye of beholder: Joint learning of gaze and actions in first person video
Yin Li, Miao Liu, and James M Rehg. 2018 · 2018
Earlier work this paper cites.
Conversational image editing: Incremental intent identification in a new dialogue task
Ramesh Manuvinakurike, Trung Bui, Walter Chang, and Kallirroi Georgila. 2018 · 2018
Cited alongside, same era.
Charades-ego: A large-scale dataset of paired third and first person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. 2018 · 2018
Cited alongside, same era.
Higs: Hand interaction guidance system
Yao Lu and Walterio Mayol-Cuevas. 2019 · 2019
Cited alongside, same era.
Situated and interactive multimodal conversations
Seungwhan Moon, Satwik Kottur, Paul Crook, Ankita De, Shivani Poddar, Theodore Levin, David Whitney, Daniel Difranco, Ahmad Beirami, Eunjoon Cho, Rajen Subba, and Alborz Geramifard. 2020 · 2020
Cited alongside, same era.
Mixed reality guidance system for motherboard assembly using tangible augmented reality
Arvin Christopher C Reyes, Neil Patrick A Del Gallego, and Jordan Aiko P Deja. 2020 · 2020
Cited alongside, same era.
Caise: Conversational agent for image search and editing
Hyounghun Kim, Doo Soon Kim, Seunghyun Yoon, Franck Dernoncourt, Trung Bui, and Mohit Bansal. 2022 · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022 · 2022
Later among the works it cites.
Assembly101: A large-scale multi-view video dataset for understanding procedural activities
Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. 2022 · 2022
Later among the works it cites.
Merlot reserve: Neural script knowledge through vision and language and sound
Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Fine-grained egocentric hand-object segmentation: Dataset, model, and applications
Lingzhi Zhang, Shenghao Zhou, Simon Stent, and Jianbo Shi. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2020 · 2020
Cited alongside, same era.
MindCraft: Theory of mind modeling for situated dialogue in collaborative tasks
Cristian-Paul Bara, Sky CH-Wang, and Joyce Chai. 2021 · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ B. Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen Creel, Jared Quincy Davis, Dorottya Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, and et al. 2021 · 2021
Cited alongside, same era.
A survey of human-in-the-loop for machine learning
Xingjiao Wu, Luwei Xiao, Yixuan Sun, Junhang Zhang, Tianlong Ma, and Liang He. 2021 · 2021
Cited alongside, same era.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Do as i can and not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, and Andy Zeng. 2022 · 2022
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022 · 2022
Cited alongside, same era.
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, Yunxin Jiao, and Aditya Ramesh. 2023 · 2023
Closest in time.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alex Herzog, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Lisa Lee, Tsang-Wei Edward Lee, Sergey Levine, Yao Lu, Henryk Michalewski, Igor Mordatch, Karl Pertsch, Kanishka Rao, Krista Reymann, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Pierre Sermanet, Jaspiar Singh, Anikait Singh, Radu Soricut, Huong Tran, Vincent Vanhoucke, Quan Vuong, Ayzaan Wahid, Stefan Welker, Paul Wohlhart, Jialin Wu, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich. 2023 · 2023
Closest in time.
Vision-language models as success detectors
Yuqing Du, Ksenia Konyushkova, Misha Denil, Akhil Raju, Jessica Landon, Felix Hill, Nando de Freitas, and Serkan Cabi. 2023 · 2023
Closest in time.
Dall-e-bot: Introducing web-scale diffusion models to robotics
Ivan Kapelyukh, Vitalis Vosylius, and Edward Johns. 2023 · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Closest in time.
Introducing whisper
OpenAI. 2022 · 2023
Closest in time.
Gpt-4v(ision) system card
OpenAI. 2023b · 2023
Closest in time.
Kosmos-2: Grounding multimodal large language models to the world
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023 · 2023
Closest in time.
Teach: Task-driven embodied agents that chat
Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gokhan Tur, and Dilek Hakkani-Tur. 2022 · 2025
Closest in time.