Fetching the paper…
Reading the bibliography…
Intelligent assistance involves not only understanding but also action.
Implicit human computer interaction through context
Albrecht Schmidt · 2000
Earlier work this paper cites.
Interactive public ambient displays: transitioning from implicit to explicit, public to personal, interaction with multiple users
Daniel Vogel and Ravin Balakrishnan · 2004
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Personalization of context-dependent applications through trigger-action rules
Giuseppe Ghiani, Marco Manca, Fabio Paternò, and Carmen Santoro · 2017
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2018
Earlier work this paper cites.
Zero-shot user intent detection via capsule neural networks
Congying Xia, Chenwei Zhang, Xiaohui Yan, Yi Chang, and Philip S Yu · 2018
Earlier work this paper cites.
Guidelines for human-ai interaction
Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al · 2019
Earlier work this paper cites.
Egovqa - an egocentric video question answering benchmark dataset
Chenyou Fan · 2019
Earlier work this paper cites.
The epic-kitchens dataset: Collection, challenges and baselines
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2020
Earlier work this paper cites.
Platform for situated intelligence
Dan Bohus, Sean Andrist, Ashley Feniello, Nick Saw, Mihai Jalobeanu, Patrick Sweeney, Anne Loomis Thompson, and Eric Horvitz · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
Enabling detailed action recognition evaluation through video dataset augmentation
Jihoon Chung, Yu Wu, and Olga Russakovsky · 2022
Earlier work this paper cites.
Actionsense: A multimodal dataset and recording framework for human activities using wearable sensors in a kitchen environment
Joseph DelPreto, Chao Liu, Yiyue Luo, Michael Foshey, Yunzhu Li, Antonio Torralba, Wojciech Matusik, and Daniela Rus · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Earlier work this paper cites.
Large language models can self-improve, 2022
Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han · 2022
Earlier work this paper cites.
Egotaskqa: Understanding human tasks in egocentric videos
Baoxiong Jia, Ting Lei, Song-Chun Zhu, and Siyuan Huang · 2022
Earlier work this paper cites.
Exploring spatial ui transition mechanisms with head-worn augmented reality
Feiyu Lu and Yan Xu · 2022
Earlier work this paper cites.
Conflab: A data collection concept, dataset, and benchmark for machine analysis of free-standing social interactions in the wild
Chirag Raman, Jose Vargas Quiros, Stephanie Tan, Ashraful Islam, Ekin Gedik, and Hayley Hung · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Cited alongside, same era.
Webshop: Towards scalable real-world web interaction with grounded language agents
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan · 2022
Cited alongside, same era.
Look ma, no hands! agent-environment factorization of egocentric videos
Matthew Chang, Aditya Prakash, and Saurabh Gupta · 2023
Cited alongside, same era.
Opening the vocabulary of egocentric actions
Dibyadip Chatterjee, Fadime Sener, Shugao Ma, and Angela Yao · 2023
Cited alongside, same era.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al · 2023
Cited alongside, same era.
Bitnet: Scaling 1-bit transformers for large language models, 2023
Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Huaijie Wang, Lingxiao Ma, Fan Yang, Ruiping Wang, Yi Wu, and Furu Wei · 2023
Later among the works it cites.
Ego-only: Egocentric action detection without exocentric transferring
Huiyu Wang, Mitesh Kumar Singh, and Lorenzo Torresani · 2023
Later among the works it cites.
Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world
Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, et al · 2023
Later among the works it cites.
Xair: A framework of explainable ai in augmented reality
Xuhai Xu, Anna Yu, Tanya R Jonker, Kashyap Todi, Feiyu Lu, Xun Qian, João Marcelo Evangelista Belo, Tianyi Wang, Michelle Li, Aran Mun, et al · 2023
Later among the works it cites.
Sigma: An open-source interactive system for mixed-reality task assistance research
Dan Bohus, Sean Andrist, Nick Saw, Ann Paradiso, Ishani Chakraborty, and Mahdi Rad · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Cited alongside, same era.
Cogagent: A visual language model for gui agents
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, et al · 2023
Cited alongside, same era.
Instruct2act: Mapping multi-modality instructions to robotic actions with large language model
Siyuan Huang, Zhengkai Jiang, Hao Dong, Yu Qiao, Peng Gao, and Hongsheng Li · 2023
Cited alongside, same era.
Towards mitigating llm hallucination via self reflection
Ziwei Ji, Tiezheng Yu, Yan Xu, Nayeon Lee, Etsuko Ishii, and Pascale Fung · 2023
Cited alongside, same era.
Single-stage visual query localization in egocentric videos
Hanwen Jiang, Santhosh Kumar Ramakrishnan, and Kristen Grauman · 2023
Cited alongside, same era.
Jiarun Liu, Wentao Hu, and Chunhong Zhang · 2023
Cited alongside, same era.
Large language model guided tree-of-thought
Jieyi Long · 2023
Cited alongside, same era.
Closest in time.
Recurrentgemma: Moving past transformers for efficient open language models
Aleksandar Botev, Soham De, Samuel L Smith, Anushan Fernando, George-Cristian Muraru, Ruba Haroun, Leonard Berrada, Razvan Pascanu, Pier Giuseppe Sessa, Robert Dadashi, et al · 2024
Closest in time.
Tri Dao and Albert Gu · 2024
Closest in time.
Augmented object intelligence: Making the analog world interactable with xr-objects
Mustafa Doga Dogan, Eric J Gonzalez, Andrea Colaco, Karan Ahuja, Ruofei Du, Johnny Lee, Mar Gonzalez-Franco, and David Kim · 2024
Closest in time.
Yifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang, Lijin Yang, Baoqi Pei, Hongjie Zhang, Lu Dong, Yali Wang, Limin Wang, et al · 2024
Closest in time.
Apple vision pro
Apple Inc · 2024
Closest in time.
Snap spectacles
Snap Inc · 2024
Closest in time.
How far can we go with synthetic user experience research?
Jie Li · 2024
Closest in time.
Human I/O: Towards a Unified Approach to Detecting Situational Impairments
XingyuBruce Liu, JiahaoNick Li, David Kim, Xiang’Anthony’ Chen, and Ruofei Du · 2024
Closest in time.
The era of 1-bit llms: All large language models are in 1.58 bits, 2024
Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, and Furu Wei · 2024
Closest in time.
Meta quest virtual reality headset
Inc. Meta Platforms · 2024
Closest in time.
Ray-ban stories smart glasses
Inc. Meta Platforms and EssilorLuxottica · 2024
Closest in time.
Ego4d goal-step: Toward hierarchical understanding of procedural activities
Yale Song, Eugene Byrne, Tushar Nagarajan, Huiyu Wang, Miguel Martin, and Lorenzo Torresani · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Closest in time.